Robots Atlas>ROBOTS ATLAS
Evaluation

BabyLM

2023ActivePublished: 25 August 2026Updated: 25 August 2026Published
Key innovation
Shifted language-model pretraining research from the scale-more-data paradigm toward sample efficiency — learning from a data budget comparable to what a child hears (10M and 100M words).
Category
Evaluation
Abstraction level
Paradigm
Operation level
TrainingData
Use cases
Sample-efficient pretrainingData-efficient pretrainingCognitively plausible language modelingCognitive modeling of language acquisitionBenchmarking small language modelsDemocratizing pretraining research

How it works

The organizers release fixed training corpora and define tracks with a hard data cap. Participants pretrain any architecture using only that budget (with a limit of up to 10 epochs) and then submit the model to a shared evaluation pipeline. The main tracks are Strict (100M words, fixed dataset), Strict-Small (10M words, fixed dataset) and Loose — a track with the same limit on text quantity but freedom in the choice of data, its domain and even its modality (which in the 2024 edition evolved into a multimodal track: 100M words + paired image-text data). Models are ranked on the BLiMP, (Super)GLUE and EWoK benchmarks.

Problem solved

Modern language models require trillions of tokens and vast compute, making pretraining research inaccessible to most labs and detached from how humans learn language from limited input. BabyLM standardizes a comparable, small data budget and a shared evaluation, enabling rigorous study of data efficiency.

Components

Strict trackCompetition track

Pretraining on a fixed, released corpus with a 100M-word budget.

Strict-Small trackCompetition track

Pretraining on a fixed, released corpus with a 10M-word budget — the most restrictive track.

Loose trackCompetition track

Only the amount of text is capped; freedom in the choice of data, domain and modality. In 2024 it evolved into a multimodal track (100M words + paired image-text data).

Multimodal trackCompetition track (since 2024)

A 100M-word text budget plus paired image-text data; evaluated on (visual) QA and grounding tasks among others. In 2024 no submission beat the baselines.

Evaluation pipelineShared evaluation

Shared evaluation pipeline: BLiMP (grammatical knowledge, minimal pairs), (Super)GLUE (language understanding after fine-tuning) and, since 2024, EWoK (Elements of World Knowledge).

Evolution

Original paper · 2023 · CoNLL 2023 · Alex Warstadt
Call for Papers — The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus
Alex Warstadt, Leshem Choshen, Aaron Mueller, Adina Williams, Ethan Wilcox, Chengxu Zhuang
2023
First edition (CoNLL 2023)
Inflection point

Inaugural challenge at CoNLL 2023 in Singapore. Strict (100M), Strict-Small (10M) and Loose tracks; evaluation on BLiMP and (Super)GLUE among others.

2024
Second edition and GPT-BERT win (CoNLL 2024)
Inflection point

Second edition at CoNLL 2024 in Miami. Added a multimodal track (100M words + image) and the EWoK benchmark. Of 31 submissions, the best was the hybrid causal-masked model GPT-BERT.

2025
Continuation as an annual workshop

The challenge continues as an annual workshop co-located with CoNLL/EMNLP, with evolving tracks and data budgets.

Hyperparameters (configurable axes)

Word budgetCritical

Hard cap on training data: 10M or 100M words.

10M wordsStrict-Small track
100M wordsStrict, Loose and Multimodal tracks
TrackHigh

Choice of competition track: Strict, Strict-Small, Loose or Multimodal.

Training epochsMedium

Training is limited to at most 10 epochs over the data budget.

≤ 10 epochsRule limit