Robots Atlas>ROBOTS ATLAS
Artificial Intelligence

Brain signals lift LLM reasoning by up to 13 percentage points

Sir Robot1 October 2026 · 3 min read
Brain signals lift LLM reasoning by up to 13 percentage points

A Peking University and Tsinghua University team showed that fMRI brain activity recorded during logic tasks can directly steer the internal representations of LLMs. The paper, published 3 August 2026 in Nature Machine Intelligence, moves the brain–model relationship from correlation to guidance, worth up to 13 percentage points of accuracy.

Key takeaways

  • Ten open models tested, 1.5B to 72B parameters — Qwen, Llama, Mistral, Phi and Gemma.
  • Signal source: the open fMRI dataset OpenNeuro ds003076 — 70 deductive problems, 10 subjects after exclusions.
  • NARI works at inference time without touching weights, flipping 100% of wrong answers across six models.
  • NARF bakes the same directions into parameters during fine-tuning and complements standard language supervision.
  • Peak gain: 13 percentage points of absolute accuracy, with transfer across reasoning types.

Reasoning does not live in the language network

The starting point is a neuroscience finding: deductive reasoning engages a Fronto-parietal network: A distributed network of frontal and parietal regions active during tasks that demand cognitive control and reasoning. largely separable from Brain language network: Left-hemisphere regions that handle language processing, largely distinct from the reasoning network.. In ds003076, participants judged a conclusion's validity after three sequential premises, and the tasks use pseudowords that strip out semantic confounds.

Alignment was measured with a neural predictivity metric built on ridge regression. Models explain a substantial share of the explainable variance in reasoning regions in aggregate, but for individual task types — syllogistic and transitive — the score drops. Convergence and divergence at once, not a simple analogy.

From measuring similarity to steering

NARI derives directions in the model's representation space from the joint structure of model activations and the fMRI: Functional magnetic resonance imaging — measures brain activity through changes in blood oxygenation across regions. signal, then edits representations at inference time without updating weights. On six models that erred on both task types, it flipped every incorrect answer to a correct one.

Baselines settle the question. Random data in place of real fMRI, and randomly sampled intervention directions, both perform clearly worse — what works is the structure of the human signal, not the model's own geometry. NARF carries those directions into parameters during fine-tuning. The approach also held on the HCP Relational Processing task and the FOLIO dataset, though effectiveness still depends on model–subject coupling.

Measurement
fMRI signal from the fronto-parietal network
Method
Deriving a direction in representation space
Do we update the model's weights?
YES
NARF — fine-tuningAllow
NO
NARI — inference-time interventionAllow
Result
Corrected model answerAllow
13percentage points — peak gain in absolute accuracyNature Machine Intelligence

Why it matters

Until now, LLM–brain comparisons stopped at observing representational similarity. Here the neural signal becomes an engineering tool — supervision that cannot be manufactured from text alone. If the effect holds, neural data will join corpora and human preferences in the training stack. It also raises an awkward question about scale: an fMRI session is expensive, and the result depends on a specific person.

What next?

  • Code is public in the pkuxmq/Brain-guided_LLM repository, so replication on other models needs no new fMRI data.
  • The authors flag model–subject coupling as the main limitation — transfer to new people is the next test.

Sources

Share this article