Robots Atlas>ROBOTS ATLAS
Artificial Intelligence

Grok 4.7: Two Index Points for 2.25x the Tokens

Sir Robot30 September 2026 · 3 min read
Grok 4.7: Two Index Points for 2.25x the Tokens

SpaceXAI released Grok 4.7 on 21 September 2026. The model lifted its Artificial Analysis index score from 44 to 46 points, but burns roughly 81,000 output tokens per task against 36,000 for its predecessor. Pricing is unchanged: $2 per million input tokens and $6 per million output.

Key takeaways

  • AA Intelligence Index: 46 points against 44 for Grok 4.6 — into the top four labs
  • Around 81,000 output tokens per task, 2.25 times more than Grok 4.6
  • Pricing unchanged: $2 per million input tokens, $6 per million output
  • Terminal-Bench 4.0: 37.6% per SpaceXAI, 26% in the standard Artificial Analysis harness
  • Context window of 500,000 tokens — the same as Grok 4.6 and 4.5

Two points higher, more than twice the tokens

Grok 4.7 is a new, larger base model with a longer reinforcement learning run weighted towards tasks that take many hours. Artificial Analysis confirms the coding gains: the Coding Agent Index rises from 47 to 56 points, DeepSWE v1.1 from 65% to 73%, and Terminal-Bench 4.0 from 18% to 33%. The discrepancy with the official figure comes from a different Harness: The layer between model and task: the call loop, the tools and the system prompts. The same model scores differently in a different harness. — SpaceXAI measures inside its own Grok Build.

81,000output tokens per task, 2.25 times more than Grok 4.6 and nearly three times more than GPT-6 AstraArtificial Analysis

There are regressions too: AA-LCR drops 3.7 percentage points, AutomationBench-AA 1.1.

…
Symbol meaning
…
output-token cost of a single task
…
price per output token ($6 per million)
…
output tokens (81,000 for Grok 4.7, 36,000 for Grok 4.6)

Musk lowered expectations a week before launch

The bar was set low by the company’s own chief before the model reached users.

Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance.

Elon Musk, founder of SpaceXAI, posting on 14 September 2026.

After launch he conceded that on agentic coding the company sits third — behind Anthropic and OpenAI.

The price advantage lasted one day

At $2 and $6 the model was the cheapest in the frontier class. The next day Anthropic shipped Claude Opus 5.5 and OpenAI shipped GPT-6 Sol and Luna, halving prices — GPT-6 Sol now costs exactly the same on input. Grok 4.7 is left with 46 points, level with MiMo-V2.6-Pro from Xiaomi.

ModelInputOutputAA Index
Claude Opus 5.5$4$2058
Claude Fable 5.1$10$5053
GPT-6 Astra$10$5053
GPT-6 Sol$2$1048
Grok 4.7$2$646
GPT-6 Luna$0.10$0.50—

Grok 4.7 API parameters:

grok-4.7model identifier
low / medium / high / xhighreasoning effort, high by default
500,000 tokenscontext window, unchanged from 4.6
200,000 tokenslong-context threshold, above which rates double

Why it matters

A token is a billing unit, not a measure of quality. A model that is cheaper per thousand tokens can still cost more on a finished task if it needs twice as many — and that is exactly what shows up here. For teams costing an agent’s work rather than reading a rate card, an advertised cut can be fiction. A price war only means something once the comparison is the bill for a completed task.

What’s next

  • Musk has announced Grok 4.8 as a noticeable improvement and Grok 4.9 as an Astra/Fable-class model
  • The Grok 4.7 Fast variant stays out of the public API — it is available only in Cursor and Grok Build
  • Parameter count and training-data composition have not been disclosed by the company

Sources

Share this article