Robots Atlas>ROBOTS ATLAS
Data

C4

2019ActivePublished: 29 September 2026Updated: 29 September 2026Published
Key innovation
A massive, cleaned text corpus from Common Crawl (~750 GB), built for the T5 model — one of the foundations of modern LLM pretraining.
Category
Data
Abstraction level
Primitive
Operation level
Data
Use cases
LLM pretraining (T5 and successors)A reference point for text corporaResearch on web-data quality and biasA base corpus for variants (mC4)

How it works

The cleaning pipeline applies heuristics: keep only lines ending with punctuation, drop pages hitting a "bad words" list, deduplicate three-sentence spans, detect language (English), remove code/placeholders and too-short documents. The result was released via TensorFlow Datasets, enabling reproducible pretraining.

Problem solved

Raw Common Crawl is huge but full of junk, duplicates and low-quality content. C4 provides a cleaned, reproducible corpus for large-scale pretraining.

Components

Common Crawl snapshotInput

Raw source of web pages.

Cleaning heuristicsProcessing

Language, quality, dedup and boilerplate filters.