Robots Atlas>ROBOTS ATLAS
Training

Social Learning

ActivePublished: 25 August 2026Updated: 25 August 2026Published
Key innovation
Ports observational, imitation-based learning from psychology into AI: an agent acquires skills and knowledge from other agents or humans, rather than solely from its own trial-and-error experience or a static data corpus.
Category
Training
Abstraction level
Paradigm
Operation level
TrainingAgent runtimeSystem
Use cases
Multi-agent reinforcement learning (multi-agent RL)Embodied AI and robotics: learning from demonstrationCultural transmission of skills between agentsImitation learningPrivacy-conscious knowledge sharing between LLMsAccelerating exploration and peer-based curricula

How it works

The mechanism depends on the variant. In multi-agent RL and robotics, an observer agent perceives the behaviour or demonstrations of another (expert) agent and updates its own policy to reproduce them โ€” via imitation learning, behavioural cloning or rewards for following the expert; in the cultural-transmission approach this happens in real time and few-shot, without pre-collected human data. In the LLM variant, a teacher model encodes task knowledge as natural-language prompts or synthetic examples that a student model uses to learn the task, limiting memorisation of raw data. Across variants the shared core is a social signal channel (observation, demonstration or message) and an update rule that transfers the observed competence to the learner.

Problem solved

Solo reinforcement learning is sample-inefficient โ€” an agent must discover good behaviours by trial and error, which is slow and can be unsafe in the real world. Social learning shortens this process by letting an agent inherit knowledge already acquired by an expert or by other agents. In the LLM setting it additionally addresses transferring knowledge between models without sharing raw, potentially sensitive training data.

Implementation

Implementation pitfalls
Imitating a suboptimal expertHigh

The learner copies the demonstrator's mistakes or biases; the learner's policy quality is bounded by the quality of the social source.

Fix:Filter and weight demonstrations by expert competence; combine with the agent's own reward signal.
Over-imitationMedium

The agent reproduces irrelevant or incidental elements of the expert's behaviour rather than only the causally relevant steps.

Fix:Learn causal task features and use mechanisms that down-weight imitation as competence grows.
Privacy leakage in cross-model knowledge transferMedium

In the LLM variant, synthetic examples or prompts may inadvertently reveal the teacher's raw training data.

Fix:Control memorisation, filter generated content, and prefer abstract prompts over verbatim examples.

Evolution

Original paper ยท 2022 ยท arXiv (DeepMind) ยท Cultural General Intelligence Team (DeepMind)
Learning Robust Real-Time Cultural Transmission without Human Data
Cultural General Intelligence Team (DeepMind), Avishkar Bhoopchand, Edward Hughes
1977
Bandura's social learning theory
Inflection point

Albert Bandura formalises learning through observation, imitation and modelling, including vicarious reinforcement โ€” the psychological foundation for later AI framings.

2022
Real-time cultural transmission in RL without human data (DeepMind)
Inflection point

DeepMind's Cultural General Intelligence Team demonstrates agents that learn new tasks in real time by imitating an expert, without pre-collected training data.

2023
Social learning for LLMs (Mohtashami et al.)

A framework is proposed in which language models transfer knowledge to one another in natural language via prompts or synthetic examples, in a privacy-conscious manner.