Robots Atlas>ROBOTS ATLAS
Training

DAgger

2011ActivePublished: 28 September 2026Updated: 28 September 2026Published
Key innovation
Recast imitation learning as a no-regret online learning problem: iteratively query the expert for correct actions on states the learner actually visits and aggregate that data, reducing compounding error from quadratic to linear in the task horizon.
Category
Training
Abstraction level
Pattern
Operation level
TrainingRobot control
Use cases
Learning robot control policies via imitationAutonomous driving and vehicle control (navigation from demonstrations)Control from operator demonstrations (teleoperation → autonomy)Structured prediction / sequence labelingManipulation and locomotion in robotics

How it works

Step by step: (1) Initialization — collect expert demonstrations and train an initial policy (the first iteration with β1=1 is ordinary behavior cloning). (2) Mixed-policy rollout — at iteration i, act with πi = βi·π* + (1−βi)·π̂i, blending the expert π* with the learned policy π̂i; the coefficient βi decays over iterations (e.g. βi = p^(i−1)) until the learner acts on its own. (3) Labeling — for the states visited by πi, query the expert for the correct actions π*(s). (4) Aggregation — append the new (s, π*(s)) pairs to the aggregate dataset D = D ∪ Di. (5) Retraining — train a new policy π̂i+1 on the entire D via supervised learning. (6) Repeat for N iterations and pick the best policy (e.g. by validation). The no-regret framing guarantees a good policy exists in the sequence and that error grows linearly in the horizon T.

Problem solved

Naive behavior cloning suffers from distribution shift: a model trained on expert states does not know the states it reaches after its own mistakes, so errors compound quadratically in the task horizon. DAgger closes this gap by training the policy on the states it actually visits.

Components

Expert (oracle)Provides correct actions for states the learner visits

A reference policy (human, MPC controller, planner) queried for π*(s) in the states encountered by the mixed policy. Its availability to label new states is a core requirement of DAgger.

Official

Mixed policyGenerates the state distribution for data collection in a given iteration

A combination of expert and learned policy controlled by βi, decaying from 1 (pure expert, behavior cloning) to 0 (fully autonomous learner), gradually shifting the collected state distribution toward the learner’s states.

Aggregated datasetAccumulates (state, expert action) pairs across all iterations

A growing set of all (s, π*(s)) pairs collected so far. Training on the full D (not just the newest batch) stabilizes learning and gives the method its name, "Dataset Aggregation".

Supervised learnerTrains the new policy on the aggregated dataset

Any supervised classifier/regressor (neural network, forest, SVM) that learns the state→action mapping on D. DAgger treats it as a no-regret online learner, which yields the theoretical guarantees.

Official

Implementation

Implementation pitfalls
Expert query costHigh

DAgger requires an expert able to label ANY state the learner visits, including erroneous ones. For a human expert this is costly and sometimes infeasible (labeling out of trajectory context).

Fix:Use an automated expert (MPC, planner) or query-reducing variants (e.g. DAgger by Coaching, SafeDAgger, HG-DAgger).
Choosing the β scheduleMedium

Decaying β too fast pushes the learner into dangerous states before it is ready; too slow keeps data in the expert’s distribution and preserves the distribution shift.

Fix:Start from β1=1 and decay exponentially; for safety-critical settings consider SafeDAgger/HG-DAgger with expert intervention.
Growing cost of training on DLow

Because D grows every iteration, the cost of retraining from scratch on the full D grows linearly with the number of iterations.

Fix:Incremental training / warm-start from the previous policy, or subsampling/weighting of older data.

Evolution

Original paper · 2011 · AISTATS 2011 · Stéphane Ross
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
Stéphane Ross, Geoffrey J. Gordon, J. Andrew Bagnell
2010
arXiv preprint published

Ross, Gordon and Bagnell publish the preprint (arXiv:1011.0686) introducing DAgger and the reduction of imitation learning to online learning.

2011
Presented at AISTATS 2011
Inflection point

The paper appears at the 14th AISTATS; DAgger becomes a standard interactive imitation-learning algorithm with a guarantee of linear (rather than quadratic) error growth.

2017
Safe / query-reducing variants (SafeDAgger, HG-DAgger)

Extensions emerge that limit expert queries and improve safety during data collection (e.g. SafeDAgger, later HG-DAgger).

Hyperparameters (configurable axes)

β (mixing) scheduleHigh

How the mixing coefficient βi is decayed (e.g. βi = p^(i−1)). β1=1 gives behavior cloning; fast decay speeds the shift to learner states, slow decay stabilizes early iterations.

β1 = 1First iteration = pure behavior cloning
βi = p^(i−1)Exponential decay
Number of iterations NHigh

Number of data-collection and retraining rounds. More iterations cover the learner’s state distribution better at the cost of more expert queries.

Base learner (policy)Medium

The supervised model learning the state→action mapping. DAgger’s guarantees assume it is a no-regret learner.

Parallelism

Parallelism level
Sequential

DAgger is inherently sequential: each iteration needs the previous iteration’s policy to collect new states, so iterations cannot be parallelized. Only rollout collection and training within a single iteration can be parallelized.

Scope
Training

Hardware requirements

Primary

DAgger is a meta-level algorithm (a data-collection and retraining loop) independent of hardware; compute requirements come from the chosen base learner, not from DAgger itself.