Startup mimic robotics, together with Black Forest Labs, has unveiled FLUX-mimic — a video-action model for training industrial robots that cuts the time to learn a single task from over 30 hours to about 30 minutes. It already runs on an Audi line handling soft, flexible parts. The announcement came on July 27, 2026.
Key takeaways
- FLUX-mimic is built on Black Forest Labs' FLUX 3 generative video model with an added action decoder.
- Task training time: from over 30 hours to about 30 minutes, depending on difficulty.
- Deployment cycles cut from months to weeks.
- Live deployment at Audi Production Lab — manipulation of soft, flexible parts.
- It builds on the earlier mimic-video model, which introduced the video-action model concept.
A video-action model instead of a VLA
FLUX-mimic is what the company calls a Video-Action Model?Video-Action Model: Pairs a generative video model (which “understands” how motion looks in the world) with a layer that turns predicted frames into robot movement commands.. Instead of teaching a robot the physics of the world from scratch using scarce demonstrations, it takes the FLUX 3 generative video model, already trained on huge amounts of footage, which “understands” the dynamics of a scene — how objects move, bend and fall.
On that base, mimic adds an action decoder?Action decoder: A model layer that translates predicted images (upcoming video frames) into concrete movement commands for the robot's arms and gripper. that turns predicted images into concrete robot movements.
The difference from Vision-Language-Action (VLA) models is fundamental. VLAs learn physics mainly from rare and expensive robot demonstration data. FLUX-mimic inherits knowledge of dynamics from video, of which there is far more. The result: fewer robot-specific examples are needed to teach a new task.
Thirty hours to thirty minutes
The strongest number is cutting the training of a single task from over 30 hours to about 30 minutes, depending on complexity. At industrial scale that is the difference between a deployment taking months and one taking weeks.
The company stresses that the key targets are tasks that have been hard for robots: manipulating soft, deformable parts. Rigid objects are predictable, but a cable, fabric or gasket changes shape in the grip. A video model that has seen how such objects behave handles them better than a rule-based approach.
Tested on an Audi line
FLUX-mimic is not just a demo. It runs at Audi Production Lab on complex soft-body manipulation in automotive production — tasks that previously required manual work.
The deployment is a collaboration between Stephan-Daniel Gravert, co-founder and chief product officer at mimic robotics, Robin Rombach, co-founder and CEO of Black Forest Labs (the maker of the FLUX models), and Christoph Schneider from Audi Production Lab.
It extends mimic's earlier work — the mimic-video model, which introduced the video-action model concept itself. FLUX-mimic moves it onto a stronger base model and shows it working in a real factory, not just a lab.
Why it matters
The biggest bottleneck in robot learning is data. Robot demonstrations are expensive and slow to collect, and VLA models need a lot of them. Making a generative video model the foundation flips the problem: the internet holds a huge amount of footage of the world, and a model that already grasps the physics of motion needs far fewer robot examples to fine-tune.
The split of roles is also telling. Black Forest Labs supplies the base model (FLUX), mimic adds the robot-control layer. It is the same pattern seen across applied robotics — a foundation-model maker and a robotics integrator working separately. If the training-time promise holds beyond Audi, the video-action approach could become a real alternative to today's dominant VLA models.
What's next
- Results at Audi Production Lab will show whether cutting training to minutes holds across varied soft-body tasks.
- Black Forest Labs' progress on the FLUX base model will directly shape the capabilities of future FLUX-mimic versions.
- An open question is whether the video-action approach scales beyond soft-part manipulation to broader industrial tasks.
Sources
- Robotics Business News — mimic robotics Introduces FLUX-mimic for Faster Industrial Robot Learning
- mimic robotics — Official site (Video-Action Models)





