Robots Atlas>ROBOTS ATLAS
Alignment

Representation Engineering (RepE)

2023ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
A top-down approach to AI transparency and control: it analyses and steers high-level concepts encoded in a network representations rather than individual neurons.
Category
Alignment
Abstraction level
Paradigm
Operation level
InferencePost-trainingModel
Use cases
Steering LLM behaviour (truthfulness, safety)Detecting lying/hallucination via representation readingControlling concepts without weight retrainingAI-transparency research

How it works

RepE collects model activations for stimuli contrasting a concept (e.g. true vs false), derives a concept direction in representation space (Representation Reading) and uses it for steering: adding/subtracting the direction to activations at inference amplifies or suppresses the trait (Representation Control), while readings also enable detection (e.g. spotting model lying).

Problem solved

Understanding and controlling LLM behaviour at the level of individual neurons is hard and does not scale. RepE operates at the representation level, where high-level concepts are more legible and steerable.

Components

Representation ReadingAnalysis

Locating directions/concepts in representations from contrastive stimuli.

Representation ControlControl

Steering behaviour by editing activations along the concept direction.