Data
C4
2019ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key
innovation
A massive, cleaned text corpus from Common Crawl (~750 GB), built for the T5 model — one of the foundations of modern LLM pretraining.
Category
Data
Abstraction level
Primitive
Operation level
Data
Use cases
LLM pretraining (T5 and successors)A reference point for text corporaResearch on web-data quality and biasA base corpus for variants (mC4)
How it works
The cleaning pipeline applies heuristics: keep only lines ending with punctuation, drop pages hitting a "bad words" list, deduplicate three-sentence spans, detect language (English), remove code/placeholders and too-short documents. The result was released via TensorFlow Datasets, enabling reproducible pretraining.
Problem solved
Raw Common Crawl is huge but full of junk, duplicates and low-quality content. C4 provides a cleaned, reproducible corpus for large-scale pretraining.
Components
Common Crawl snapshotInput
Raw source of web pages.
Cleaning heuristicsProcessing
Language, quality, dedup and boilerplate filters.
Evolution
Original paper · 2019 · Colin Raffel
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5)
Colin Raffel, Noam Shazeer, Adam Roberts, et al.
2019
C4 introduced together with the T5 model
Inflection point2021
Documentation analyses of C4 (quality, source bias)