A team at the Weizmann Institute has presented Brain-IT, a system that reconstructs a viewed image from an fMRI scan using one hour of data from a new subject instead of forty hours. The paper was accepted at ICLR 2026, and MIT Technology Review covered it on 1 October 2026. The entry cost for visual-decoding research drops by an order of magnitude.
Key takeaways
- One hour of fMRI data from a new subject matches methods trained on full 40-hour recordings
- The decoder splits the job into low-level structure and high-level semantic features
- Around 70% of the training data was never paired with scans, per MIT Technology Review
- Eight subjects, roughly 9,000 images each, shown during scanning
- Voxel resolution went from 3 to 1 cubic millimetre — a 3 mm voxel covered some 16,000 neurons
Two branches instead of one
Earlier systems tried to derive the image from a single representation, which blurred "what is there" with "how it looks". Brain-IT separates the signals: one path predicts patch-level semantic features, the other low-level structure — layout, contours and colour. Only together do they drive the generator.
The diagram shows the order of stages, not the share of computation. Splitting into two branches is the core of the method: a single decoder previously had to guess scene content and scene geometry at the same time, which degraded both whenever the signal was weak.
The glue is the brain-interaction Transformer the system is named after. Rather than treating activity as a flat vector, the model lets groups of voxels?Voxel: The smallest volume cube a scanner can resolve. The smaller it is, the fewer neurons get averaged into a single measurement. exchange information, mirroring the fact that neighbouring cortical areas encode related parts of a scene.
One hour in the scanner, not forty
The strongest result is not reconstruction quality but calibration cost. Forty hours in a scanner is a barrier that limits studies to a handful of volunteers — usually the researchers themselves. One hour fits into a single visit.
| Parameter | Earlier methods | Brain-IT |
|---|---|---|
| Calibration per new subject | 40 hours in the scanner | 1 hour |
| Voxel resolution | 3 cubic millimetres | 1 cubic millimetre |
| Training data with no paired scan | — | around 70% |
| Study participants | — | 8 subjects, ~9,000 images each |
The shortcut also comes from data. As MIT Technology Review reports, roughly 70% of the training material consisted of images with no matching scans. The model learns the visual world separately, then learns the fit to one particular brain from a small sample.
Why it matters
Decoding images from fMRI has been a lab demonstration precisely because it demanded hours of lying in a scanner. Cutting calibration to a single session moves the topic from curiosity to research instrument — and forces a conversation about neural-data privacy before any clinical or commercial practice exists.
What's next
- The paper was accepted at ICLR 2026, with the preprint already available on arXiv
- Reconstruction still requires an fMRI scanner, so applications outside the lab remain out of reach
- The claimed 1 cubic millimetre resolution comes from MIT Technology Review's reporting, not the paper abstract





