OpenAI released GPT-6 Astra on 3 September, an agentic model that drives browsers, spreadsheets and apps itself rather than saying what to click. Greg Brockman marked the launch with the words “Welcome to the AGI era”. Independent measurements are more cautious: it dominates security tasks but trails Claude Fable 5.1 overall.
Key takeaways
- Launched 3 September 2026, in the API as gpt-6-astra — $10/M input tokens and $50/M output
- ARC-AGI-3: 99.9% on OpenAI's own harness, 62.7% on the standard one
- ExploitBench: 100% against 78.5% for GPT-5.6 Sol
- Artificial Analysis Intelligence Index: 61 points, five below Claude Fable 5.1
- Critical designation for cybersecurity under OpenAI's Preparedness Framework
An agent that clicks for itself
Astra is a computer use model?Computer use model: A model that operates a computer interface the way a person does — clicking, scrolling and typing inside existing applications instead of returning text for a human to paste.: rather than generating text to paste, it drives the browser, CRM and spreadsheet itself. The rollout started with a narrow group of organisations in the Daybreak programme and reaches ChatGPT Plus, Pro, Business and Enterprise, the API, AWS Bedrock and Azure within days. OpenAI says training consumed over 100,000 DBUs on Stargate infrastructure, its largest jump yet.
The score depends on the harness
The loudest number needs an asterisk. According to ARC Prize, Astra reaches 99.9% on ARC-AGI-3 Semi-Private for roughly $19,000 — but only with its own Provider Adapter harness?Harness: The runtime layer between a model and a benchmark: it decides what the model sees between turns, how much state carries over and how long a conversation may run. The same model scores differently on different harnesses., which preserves opaque reasoning state between requests. On the standard harness the score drops to 62.7% and the cost rises to $26,000.
| Reasoning effort | Standard harness | Provider Adapter |
|---|---|---|
| max | 62.7% · $26,098 | 98.6% · $17,332 |
| xhigh | 59.3% · $37,317 | 98.4% · $18,147 |
| high | 54.8% · $40,705 | 99.9% · $18,817 |
| medium | 38.6% · $48,090 | 98.4% · $19,285 |
| low | 17.5% · $38,166 | 98.0% · $21,298 |
| none | 35.2% · $49,791 | 96.7% · $23,457 |
A more interesting result sits underneath. The model used 51.7% fewer actions per level than the median tested human and was more economical on 96% of levels. ARC Prize stresses that saturating the benchmark is not proof of AGI.
Strong on offensive code
Astra dominates wherever security is the point — 100% on ExploitBench against 78.5% for GPT-5.6 Sol, and 99.2% within four attempts on SRE-Bench against its predecessor's 68.7%. It holds 96.3% accuracy above half a million tokens. In the general Intelligence Index, though, it stops at 61 points, losing to Fable 5.1 and Muse Spark 1.3. GDPval results were missing from the launch materials.
Why it matters
The launch shifts the centre of gravity from answer quality to finishing a task inside someone else's software. That changes pricing — what counts is the cost of a completed task, not of a token. The harness gap also shows benchmark results starting to depend on provider infrastructure rather than the model itself. Cross-lab comparison gets harder.
What's next
- Broader availability across ChatGPT tiers, AWS Bedrock and Azure is due within days of 3 September
- Claude Fable 5.1 has no published ARC-AGI-3 result, so a direct comparison of the leaders remains impossible
Sources
- VentureBeat — 'Welcome to the AGI era': OpenAI launches GPT-6 Astra
- ARC Prize — OpenAI's GPT-6 Astra on ARC-AGI-3
- Simon Willison's Weblog — GPT-6 Astra





