DAgger
How it works
Step by step: (1) Initialization — collect expert demonstrations and train an initial policy (the first iteration with β1=1 is ordinary behavior cloning). (2) Mixed-policy rollout — at iteration i, act with πi = βi·π* + (1−βi)·π̂i, blending the expert π* with the learned policy π̂i; the coefficient βi decays over iterations (e.g. βi = p^(i−1)) until the learner acts on its own. (3) Labeling — for the states visited by πi, query the expert for the correct actions π*(s). (4) Aggregation — append the new (s, π*(s)) pairs to the aggregate dataset D = D ∪ Di. (5) Retraining — train a new policy π̂i+1 on the entire D via supervised learning. (6) Repeat for N iterations and pick the best policy (e.g. by validation). The no-regret framing guarantees a good policy exists in the sequence and that error grows linearly in the horizon T.
Problem solved
Naive behavior cloning suffers from distribution shift: a model trained on expert states does not know the states it reaches after its own mistakes, so errors compound quadratically in the task horizon. DAgger closes this gap by training the policy on the states it actually visits.
Components
A reference policy (human, MPC controller, planner) queried for π*(s) in the states encountered by the mixed policy. Its availability to label new states is a core requirement of DAgger.
Official
A combination of expert and learned policy controlled by βi, decaying from 1 (pure expert, behavior cloning) to 0 (fully autonomous learner), gradually shifting the collected state distribution toward the learner’s states.
A growing set of all (s, π*(s)) pairs collected so far. Training on the full D (not just the newest batch) stabilizes learning and gives the method its name, "Dataset Aggregation".
Any supervised classifier/regressor (neural network, forest, SVM) that learns the state→action mapping on D. DAgger treats it as a no-regret online learner, which yields the theoretical guarantees.
Official
Implementation
DAgger requires an expert able to label ANY state the learner visits, including erroneous ones. For a human expert this is costly and sometimes infeasible (labeling out of trajectory context).
Decaying β too fast pushes the learner into dangerous states before it is ready; too slow keeps data in the expert’s distribution and preserves the distribution shift.
Because D grows every iteration, the cost of retraining from scratch on the full D grows linearly with the number of iterations.
Evolution
Ross, Gordon and Bagnell publish the preprint (arXiv:1011.0686) introducing DAgger and the reduction of imitation learning to online learning.
The paper appears at the 14th AISTATS; DAgger becomes a standard interactive imitation-learning algorithm with a guarantee of linear (rather than quadratic) error growth.
Extensions emerge that limit expert queries and improve safety during data collection (e.g. SafeDAgger, later HG-DAgger).
Hyperparameters (configurable axes)
How the mixing coefficient βi is decayed (e.g. βi = p^(i−1)). β1=1 gives behavior cloning; fast decay speeds the shift to learner states, slow decay stabilizes early iterations.
Number of data-collection and retraining rounds. More iterations cover the learner’s state distribution better at the cost of more expert queries.
The supervised model learning the state→action mapping. DAgger’s guarantees assume it is a no-regret learner.
Parallelism
DAgger is inherently sequential: each iteration needs the previous iteration’s policy to collect new states, so iterations cannot be parallelized. Only rollout collection and training within a single iteration can be parallelized.
Hardware requirements
DAgger is a meta-level algorithm (a data-collection and retraining loop) independent of hardware; compute requirements come from the chosen base learner, not from DAgger itself.