Robots Atlas>ROBOTS ATLAS
Data

SAYCam

2020ActivePublished: 25 August 2026Updated: 25 August 2026Published
Key innovation
The first large, longitudinal audiovisual dataset recorded from a young child's own point of view, letting researchers train and test AI models on input resembling real human developmental experience.
Category
Data
Abstraction level
Building block
Operation level
DataTraining
Use cases
Developmentally plausible learningGrounded language learningSelf-supervised visual representation learningMultimodal learning (vision + speech)Word-referent mapping / word learningTesting the poverty-of-the-stimulus hypothesisData-efficiency researchCognitive and language development research

How it works

Three children wore a lightweight head-mounted camera that recorded video and audio from their own viewpoint during everyday activities, roughly 2 hours per week for about 2.5 years (ages ~6–32 months). The recordings were processed: speech was transcribed (over 200,000 words) and the material annotated with metadata allowing search by age, location, setting and objects. The data was deposited in Databrary under access control. Researchers use the video frames and paired speech transcripts to train models (for example self-supervised vision networks or multimodal contrastive models pairing images with child-directed speech).

Problem solved

AI models are typically trained on massive curated internet corpora that bear no resemblance to the input stream a child learns from. There was a lack of realistic, longitudinal data on what a learning child actually sees and hears, needed to study what can be learned from such naturalistic, limited input.

Components

Child S recordings (SAYCam-S)Single-child data subset

The largest and most widely used subset, containing egocentric recordings from the child denoted by the initial S.

Child A recordings (SAYCam-A)Single-child data subset

Subset of egocentric recordings from the child denoted by the initial A.

Child Y recordings (SAYCam-Y)Single-child data subset

Subset of egocentric recordings from the child denoted by the initial Y.

Labeled-SEvaluation set

An evaluation subset of frames from child S, manually annotated with category labels, used for testing classification and word-referent mappings.

Speech transcriptsLinguistic data layer

Over 200,000 transcribed words of speech recorded in the child's environment, time-aligned with the video material.

Evolution

Original paper · 2020 · Open Mind · Jessica Sullivan
SAYCam: A Large, Longitudinal Audiovisual Dataset Recorded From the Infant's Perspective
Jessica Sullivan, Michelle Mei, Andrew Perfors, Erica H. Wojcik, Michael C. Frank
2020
SAYCam dataset introduced and released
Inflection point

Sullivan and colleagues describe SAYCam in the journal Open Mind and deposit the data in the Databrary repository under access control.

2020
Self-supervised vision models trained on SAYCam

Emin Orhan and colleagues show that useful visual representations can be learned self-supervised from SAYCam recordings (TC-S/A/Y/SAY models).

2024
CVCL learns language from a single child's data
Inflection point

Vong and colleagues train the CVCL multimodal contrastive model on ~61 hours of a single child's recordings (ages ~6–25 months), learning word-referent mappings and showing zero-shot generalization; published in Science.