DrivingBench, an independent project, gave four frontier language models control of a real 2022 Toyota Corolla on a parking-lot cone course. Only GPT-6 Astra finished: 134.7 metres in 5 minutes and 22 seconds. None of the four was trained for autonomous driving.
Key takeaways
- GPT-6 Astra finished the course on attempt two: 134.7 m in 5:22, 24 commands
- Claude Fable 5.1 reached 45%, Grok 4.6 11% and GPT-5.6 Sol 6%
- Attempt two burned 6.6M tokens and cost 7.74 USD
- Three MCP tools wired to openpilot and comma four drove the car
- Renaming the MCP server “DrivingBench Sandbox” cut refusals
| Model | Course completed | Distance |
|---|---|---|
| GPT-6 Astra | 100% | 134.7 m |
| Claude Fable 5.1 | 45% | not reported |
| Grok 4.6 | 11% | not reported |
| GPT-5.6 Sol | 6% | not reported |
A Corolla, a comma four and three tools
A comma four plugged into the Corolla's CAN bus?CAN bus: The vehicle's internal network, where controllers exchange messages about speed, steering and braking. over OBD-C, with control built on openpilot. Cameras and telemetry travel via a laptop to an MCP server, then to the model in a Codex, Claude Code or Cursor client.
Three tools: observe() returns frames and speed, set_motion() sets direction, steering percent, speed and duration, and stop_now() brakes. Steering torque and brake pressure are openpilot's job, not the model's.
In-context learning, no weight updates
Three attempts ran in one chat. Astra's first ended at 49% and 67.3 metres, when the operator braked. A reflection prompt followed.
I declared the car aligned too early. […] I straightened and increased the requested speed to 1.5 m/s.
GPT-6 Astra, reflection quoted in the DrivingBench report.
On attempt two it never exceeded 0.8 m/s and used 100% steering on 20 of 24 commands. Average speed: 0.42 m/s, or 1.5 km/h.
What the test does not show
Each model got one session, and its three attempts shared one context. Hard limits guarded the lot: 0.5–3.5 m/s, an emergency cut-off at 6 m/s and an operator over the brake.
The forward camera cannot see obstacles right beside the car. Astra most often refused to drive on safety grounds.
Why this matters
The result says nothing about driving quality, because 1.5 km/h is walking pace. It does say a general-purpose model can close the perception–plan–control loop in the real world through an ordinary tool API and fix its own mistake without a weight update. For carmakers the point is that a car's decision layer need not come from a classic autonomous-driving stack, but from a general model driven by tools.
What's next
- DrivingBench v2 is to add multiple sessions, different reasoning efforts and a harder course
- The harness and prompts are on GitHub, but reproduction needs a Corolla and comma four
- The authors promise a separate piece on safety and alignment
Sources
- DrivingBench — Report
- DrivingBench — Leaderboard
- TMTPost — 不用专门训练自动驾驶,GPT-6 Astra已经能开真车了?
- GitHub — drivingbench_harness_v1





