Three children wore a lightweight head-mounted camera that recorded video and audio from their own viewpoint during everyday activities, roughly 2 hours per week for about 2.5 years (ages ~6–32 months). The recordings were processed: speech was transcribed (over 200,000 words) and the material annotated with metadata allowing search by age, location, setting and objects. The data was deposited in Databrary under access control. Researchers use the video frames and paired speech transcripts to train models (for example self-supervised vision networks or multimodal contrastive models pairing images with child-directed speech).
AI models are typically trained on massive curated internet corpora that bear no resemblance to the input stream a child learns from. There was a lack of realistic, longitudinal data on what a learning child actually sees and hears, needed to study what can be learned from such naturalistic, limited input.
The largest and most widely used subset, containing egocentric recordings from the child denoted by the initial S.
Subset of egocentric recordings from the child denoted by the initial A.
Subset of egocentric recordings from the child denoted by the initial Y.
An evaluation subset of frames from child S, manually annotated with category labels, used for testing classification and word-referent mappings.
Over 200,000 transcribed words of speech recorded in the child's environment, time-aligned with the video material.
Sullivan and colleagues describe SAYCam in the journal Open Mind and deposit the data in the Databrary repository under access control.
Emin Orhan and colleagues show that useful visual representations can be learned self-supervised from SAYCam recordings (TC-S/A/Y/SAY models).
Vong and colleagues train the CVCL multimodal contrastive model on ~61 hours of a single child's recordings (ages ~6–25 months), learning word-referent mappings and showing zero-shot generalization; published in Science.