Alignment
Representation Engineering (RepE)
2023ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key
innovation
A top-down approach to AI transparency and control: it analyses and steers high-level concepts encoded in a network representations rather than individual neurons.
Category
Alignment
Abstraction level
Paradigm
Operation level
InferencePost-trainingModel
Use cases
Steering LLM behaviour (truthfulness, safety)Detecting lying/hallucination via representation readingControlling concepts without weight retrainingAI-transparency research
How it works
RepE collects model activations for stimuli contrasting a concept (e.g. true vs false), derives a concept direction in representation space (Representation Reading) and uses it for steering: adding/subtracting the direction to activations at inference amplifies or suppresses the trait (Representation Control), while readings also enable detection (e.g. spotting model lying).
Problem solved
Understanding and controlling LLM behaviour at the level of individual neurons is hard and does not scale. RepE operates at the representation level, where high-level concepts are more legible and steerable.
Components
Representation ReadingAnalysis
Locating directions/concepts in representations from contrastive stimuli.
Representation ControlControl
Steering behaviour by editing activations along the concept direction.
Evolution
Original paper · 2023 · Andy Zou
Representation Engineering: A Top-Down Approach to AI Transparency
Andy Zou, et al.
2023
Zou et al. introduce Representation Engineering
Inflection point