Robocurve, an independent robotics evaluator backed by Y Combinator, wired GPT-6 Astra directly into a bimanual robot — with no learned model anywhere in the control path. On a block-moving task the model completed 19 of 20 attempts against 8 for Claude Fable 5.1. On a precision task both systems stalled at an identical score.
Key takeaways
- Block into bowl: GPT-6 Astra 19/20, Claude Fable 5.1 8/20, Claude Fable 5 1/20
- Puzzle piece into groove: Astra and Fable 5.1 both at 2/20 — the lead disappears
- Output token use: 2,100 per run against 12,900 for Fable 5.1
- $0.94 per run against $2.12, and 2.5 minutes against 6.8
- The RoboDojo team halted physical-robot testing after equipment was damaged
A control path with no learned policy
The model receives three camera views and returns absolute gripper poses. The robot’s inverse kinematics converts them into joint angles — no learned policy in between.
The diagram shows what sits between a camera frame and arm motion in the Robocurve test. There is no trained policy between model and robot — gripper poses go straight into a classical inverse-kinematics solver.
The lead vanishes where precision matters
Moving a red block into a bowl, Astra finishes 19 of 20 runs. Inserting a round puzzle piece into a matching groove, it drops to 2 of 20, exactly matching Fable 5.1, and scores worse on mean stage progress (2.00 against 2.35). The report notes that the model reaches the groove and stalls at the same final step as its rival.
| Measure | GPT-6 Astra | Claude Fable 5.1 | Claude Fable 5 |
|---|---|---|---|
| Block into bowl | 19/20 | 8/20 | 1/20 |
| Puzzle into groove | 2/20 | 2/20 | — |
| Mean stage progress (puzzle) | 2.00 | 2.35 | — |
| Output tokens per run | 2,100 | 12,900 | — |
| Cost per run | $0.94 | $2.12 | — |
| Time per run | 2.5 min | 6.8 min | — |
Physical-robot testing had to be stopped
The RoboDojo benchmark team ran a separate evaluation.
Astra frequently issued physically unreasonable or unsafe actions during real-robot trials, and some incidents damaged equipment. Testing was therefore stopped, and the complete 18-task official protocol was not run.
The RoboDojo team, report of 16 September 2026.
In simulation the picture is bimodal: across 42 tasks the model clears 50% success on ten of them, yet scores zero on sixteen.
What actually works
Online adaptation performs surprisingly well — the model detects and corrects a flipped image, inverted axis signs or random disturbance forces on its own. In-context learning runs the other way: a single demonstration cut success from 22.9% to 17.9%, and a text-described demonstration to 12.9%.
Why it matters
The block result is easy to sell as proof that a general model is enough to drive a robot. Splitting it across two tasks shows where the boundary runs: wherever contact, force and geometry decide the outcome, the advantage evaporates. Faster token generation will solve latency, but it will not supply the physical knowledge a model has no way of extracting from text and images alone.
What’s next
- Robocurve projects real-time robot control by Fable-class models somewhere between late 2026 and 2029
- The full 18-task RoboDojo-Real protocol for Astra remains unfinished
- The Robocurve comparison carries caveats: the two models’ runs were two days apart, and the block task ran on different rigs





