Robots Atlas>ROBOTS ATLAS
Training

Model Collapse

2024ActiveUpdated: 19 August 2026Published
Key innovation
Identifying and formally describing the degeneration of generative models recursively trained on data produced by earlier models — losing the tails of the distribution and converging to impoverished outputs.
Category
Training
Abstraction level
Pattern
Operation level
TrainingData
Use cases
Assessing the risk of training on AI-generated dataDesigning data-curation strategies (mixing with human data)Justifying data provenance taggingResearch on training-data quality and diversity

How it works

When a model is trained on samples generated by a previous model (rather than on original data), errors compound: finite sampling misses rare events, the model approximates the distribution imperfectly, and the learning process adds its own error. As a result each successive generation increasingly 'clips' the distribution's tails and narrows output diversity, until the model loses the ability to represent the original data diversity. The phenomenon was observed for LLMs, VAEs and Gaussian mixture models.

Problem solved

It surfaces the hidden risk of training models on synthetic/AI-generated data and explains why quality can degrade across generations — enabling mitigation strategies (human data, provenance, data mixing).

Key mechanisms

Recursive training on data from previous models
Compounding of errors: sampling, approximation, learning
Loss of distribution tails (early collapse)
Convergence to a low-variance distribution (late collapse)

Strengths & limitations

Strengths
✓Explains a measurable degradation mechanism
✓Applies across generative model classes (LLMs, VAEs, GMMs)
✓Motivates data protection and curation
Limitations
✗It is a phenomenon/risk, not a method — it needs mitigations
✗The effect size depends on the share of synthetic data and the procedure
✗It can be partly avoided by mixing with real data

Components

Statistical sampling errorDegradation source

Finite samples miss rare events from the distribution tails.

Functional approximation errorDegradation source

The model does not represent the distribution exactly.

Learning errorDegradation source

The optimization procedure introduces additional bias.

Implementation

Implementation pitfalls
Invisible quality driftHigh

Collapse can build slowly and evade shallow evaluation.

Fix:Monitor diversity/tail coverage, not just average metrics.
Training-data contaminationHigh

Web data contains an increasing share of AI-generated content.

Fix:Data provenance, filtering, a share of human data.

Evolution

Original paper · 2024 · Nature (2024) · Ilia Shumailov
AI models collapse when trained on recursively generated data
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, Yarin Gal
2023
Early work on recursion and synthetic data

Research flags the risk of training on model outputs.

2024
Nature publication (Shumailov et al.)
Inflection point

Formal description and the name 'model collapse'; early vs late collapse.

Hyperparameters (configurable axes)

Synthetic data ratioCritical
wysokiMore model-generated data → faster collapse.
z domieszką danych realnychMixing with human data slows collapse.
Number of generationsHigh
wiele iteracjiThe effect compounds with each generation.