Robots Atlas>ROBOTS ATLAS
Training

Sample-Efficient Pretraining

2023ResearchPublished: 25 August 2026Updated: 25 August 2026Published
Key innovation
Shifts the goal of pretraining from maximizing data scale to maximizing model quality under a data budget comparable to a child's language exposure (~10–100M words).
Category
Training
Abstraction level
Paradigm
Operation level
TrainingData
Use cases
Training language models on an academic budgetResearch on language acquisition and cognitive plausibility of modelsLow-resource language modelingBenchmarking data efficiency (BabyLM Challenge)Prototyping architectures and training objectives without large clusters

How it works

Instead of enlarging the training set, researchers work on extracting more learning signal from each example. Common levers are: (1) richer training objectives — e.g. a hybrid of masked (BERT) and causal (GPT) language modeling, as in GPT-BERT; (2) curriculum learning — ordering data from easier to harder; (3) data augmentation and careful preprocessing; (4) architectures optimized for small corpora (e.g. LTG-BERT); (5) regularization and multiple passes over the data without overfitting; (6) incorporating multimodal signal. Models are evaluated on standardized benchmarks (e.g. BLiMP, (Super)GLUE, EWoK in the BabyLM pipeline) to compare quality under an equal, constrained data budget.

Problem solved

Modern LLMs achieve high quality mainly by scaling data to hundreds of billions or trillions of tokens, which is expensive, inaccessible to most teams, and implausible as a model of human language acquisition. Sample-Efficient Pretraining addresses how to reach good quality with data on the scale of a child's language exposure (10–100M words).

Components

Training objectivesIncreasing learning signal per example

Richer loss functions that extract more signal per example, e.g. combining masked (MLM) and causal (CLM) language modeling.

Curriculum learningImproving learning efficiency via data ordering

Ordering training data from easier to harder, inspired by a child's exposure order.

Data augmentationIncreasing effective data diversity

Creating additional data variants and careful preprocessing to increase effective coverage under a fixed word budget.

Architecture selectionMatching model capacity to the data budget

Architectures optimized for small corpora, e.g. LTG-BERT or the hybrid GPT-BERT.

Regularization and multi-epoch trainingMaximizing use of limited data

Stronger regularization and multiple passes over the same set without overfitting, to fully exploit a small corpus.

Multimodal dataSupplementing the textual signal

Using signal from other modalities (e.g. vision) as additional context when text is limited.

Implementation

Implementation pitfalls
Overfitting on a small corpusHigh

With 10–100M words, large models easily overfit the data.

Fix:Stronger regularization, model-size selection, augmentation, early stopping.
Evaluation validityMedium

Without standardized benchmarks it is hard to fairly compare quality at an equal data budget.

Fix:Use a shared pipeline (e.g. BabyLM: BLiMP, (Super)GLUE, EWoK).
Violating the data budgetMedium

Inadvertently introducing extra data (e.g. via pretrained components) invalidates the comparison.

Fix:Strictly account for every token; train from scratch within the track.
Curriculum brittlenessMedium

Gains from curriculum learning can be unstable and task-dependent.

Fix:Validate ordering variants; do not assume universal improvement.

Evolution

Original paper · 2023 · CoNLL 2023 (BabyLM Challenge) · Alex Warstadt
Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
Alex Warstadt, Aaron Mueller, Leshem Choshen, Ethan Wilcox, Chengxu Zhuang, Juan Ciro, Rafael Mosquera, Bhargavi Paranjape, Adina Williams, Tal Linzen, Ryan Cotterell
2023
First BabyLM Challenge
Inflection point

Defined strict data tracks: Strict-Small (10M words) and Strict (100M words); established the sample-efficient pretraining paradigm at CoNLL.

2024
GPT-BERT wins BabyLM 2024
Inflection point

A hybrid of masked and causal language modeling outperforms MLM-only and CLM-only variants under the same data budget.

2026
Fourth BabyLM edition at EMNLP 2026

The challenge continues as a workshop, maintaining the focus on sample-efficient pretraining on developmentally plausible corpora.

Hyperparameters (configurable axes)

Data budgetCritical

Size of the training corpus, e.g. 10M words (Strict-Small) or 100M words (Strict).

Objective mixtureHigh

Ratio between masked and causal language modeling (e.g. in GPT-BERT).

Model sizeHigh

Parameter count chosen so as not to overfit the small corpus.

Curriculum orderMedium

How data is ordered from easy to hard.

Number of epochsMedium

How many times the model passes over the same small dataset.

Parallelism

Parallelism level
Fully parallel

The paradigm inherits transformer training parallelism; the bottleneck is the data budget, not compute parallelism.

Scope
TrainingAcross tokensAcross devices

Hardware requirements

Good fit

The small data budget enables training on a modest single GPU in academic settings.

Primary

Training the underlying transformer still benefits from GPU accelerators.