Robots Atlas>ROBOTS ATLAS
Artificial Intelligence

Skild AI Unveils S1: A Robot Learns a Task From One Video

Sir Robot3 September 2026 · 3 min read
Skild AI Unveils S1: A Robot Learns a Task From One Video

Skild AI has released S1, its flagship robot foundation model, which picks up a new task from a single video demonstration without any weight updates. The company described the model on its blog in August 2026, and on 31 August co-founder and CEO Deepak Pathak laid out the details to The Robot Report. It is an attempt to carry in-context learning from language models straight into robot control.

Key takeaways

  • S1 performs previously unseen tasks lasting up to 10 minutes after a single video demonstration
  • 66% success on unseen tasks versus 9% for language-prompted policies
  • One in-context example is worth roughly 380 post-training episodes
  • Pretraining blends teleoperation, UMI glove data, egocentric video and simulation
  • Skild AI has raised close to $1.7 billion since 2023

One video instead of fine-tuning

A new task normally means a new round of post-training: An extra learning stage after pretraining that tunes a model to one specific task.. S1 skips that step. A clip of a human performing the task goes into the prompt, and the model reproduces the demonstrator’s intent on the robot.

Skild reports 96% success on tasks seen during training at 100,000 hours of pretraining and 66% on unseen tasks — against 9% for comparable policies driven by a text description.

380post-training episodes replaced by a single video demonstration placed in contextSkild AI, S1 model card
You just add a video of a human doing something to the prompt, and the robot follows it. The tasks we are showing are complex and long — not three- or four-second trivia.

Deepak Pathak, co-founder and CEO of Skild AI.

Four data sources instead of one

S1’s pretraining mixes four data streams rather than leaning on one. Each source has a different profile: teleoperation delivers the highest quality but scales worst, egocentric video the other way round, with UMI gloves in between. The company says it spends about three dollars on quality control for every dollar spent collecting data.

The four streams behind S1’s pretraining:

teleoperationa human drives the robot directly — highest quality, worst scalability
UMI glovesa human performs the task wearing gloves that capture hand motion
egocentric videofirst-person footage — the stream that scales best
simulationa virtual environment producing data without a physical robot

Omni-bodied, but humanoids come later

The model runs on quadrupeds, static arms and humanoids — Skild calls this an omni-bodied approach. Demonstrated tasks include repotting a plant, brewing pour-over coffee and cooking pancakes. We covered a similar one-demo learning mechanism with GEN-1.5.

Why it matters

Robotics has been stuck in a model where every new task costs a separate round of data collection and training. If Skild’s numbers hold up outside controlled demos, deployment cost stops scaling linearly with the number of tasks — and that shifts the economics of the sector more than another benchmark point ever could. Pathak himself tempers expectations: this is not robotics’ ChatGPT moment yet, only the first sign of one.

What's next?

  • A demonstration of S1 in production use is promised within weeks (Pathak, 31 August 2026)
  • Warehouse deployments will build on Zebra Technologies' robotics division, formerly Fetch Robotics, acquired in April 2026
  • Scaling the model to humanoids is named by the company as the next line of work

Sources

Share this article