Robots Atlas>ROBOTS ATLAS
Grok 4.7

Grok 4.7

Grok 4.7 (grok-4.7) · Family: Grok
xAI's most capable model for coding and knowledge work (September 2026): a larger base model, longer RL training, a 500K-token context window and $2/$6 per million tokens.
✓ Active✓ Public accessLLMMultimodalReasoning modelTool-using model📁 Grok
Context window
500k
tokens
Release date
21 September 2026
Access:APIHostedDeployment:☁ Cloud

Overview

Grok 4.7 is an xAI language model released on 21 September 2026 and positioned by the company as its most capable model for coding and knowledge work. It is built on a larger base model trained with longer reinforcement learning; xAI emphasizes that the model works longer on difficult tasks and checks its own work more carefully. Document creation and multi-hour professional tasks improved noticeably — across legal, electrical engineering and office work.

The model runs under the API identifier grok-4.7, supports a 500,000-token context window and text plus image input, and its knowledge extends through May 2026. Base pricing is $2 per million input tokens and $6 per million output tokens; above 200,000 prompt tokens the rates rise to $4 and $12. It is available through the Grok API, Grok Build, Cursor and third-party platforms and agentic harnesses.

In figures reported by xAI, Grok 4.7 scores 64.0% on EEBench (electrical engineering, up from 53.0%), 46.3% on CursorBench 4.0 (up from 40.4%), 38.0% on Terminal-Bench 4.0 (up from 20.3%) and 19.6% on Harvey Legal Agent (up from 15.8%). On safety tests the model allows through 3.3% of risky cybersecurity prompts while keeping refusals low for legitimate use, and it reaches 62.4% on a biosafety benchmark.

Classification
LLMMultimodalReasoning modelTool-using model
Family: Grok
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📏 Context: 500k
Tools
📥 Input: text, image, documents

Technical specification

Context window
500k
tokens
Knowledge cutoff
1 May 2026
Knowledge boundary
Features:Tool use
Modalities
⬇ Input
textimagedocuments
⬆ Output
textcodestructured_data

Capabilities and applications

Native model capabilities
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Advanced reasoning
The ability to perform multi-step, structured reasoning: analysing problems, planning steps, and drawing conclusions from hypotheses. Reasoning-first models (e.g. GPT-5.1 Thinking) dedicate a portion of inference to chains of thought before responding.
Category: reasoning
Multi-step reasoning
Carrying out multi-step chains of reasoning across long, complex tasks.
Category: reasoning
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Agentic coding
Multi-hour, multi-step programming tasks performed autonomously by the model: cloning a repository, running tests, iterating on fixes, integrating with CLI tools. Characteristic of Codex variants (GPT-5.1-Codex-Mini, Codex-Max).
Category: coding
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Function Calling
Category: planning
Structured output
Producing data in structured formats such as JSON.
Category: structured_generation
Image understanding
Analysing and interpreting the content of images.
Category: vision
Web browsing
Ability of the model to autonomously search and browse web pages to retrieve up-to-date information.
Category: other
Multi-step project execution
The ability to autonomously drive multi-hour, multi-step projects: decomposing the task, planning the sequence of actions, iteratively delivering results, and adjusting based on feedback. Key for enterprise and knowledge-work agents.
Category: planning
Planning
Forming and executing action plans for complex tasks.
Category: planning

Benchmark results

5 benchmarks
EEBench
accuracy
64.0%%
📄 xAI — ogłoszenie Grok 4.7
Electrical engineering benchmark. Grok 4.7 scores 64.0% versus 53.0% for the previous version.
CursorBench 4.0
accuracy
46.3%%
📄 xAI — ogłoszenie Grok 4.7
Agentic coding benchmark in Cursor. Grok 4.7 scores 46.3% versus 40.4%.
Terminal-Bench 4.0
accuracy
38.0%%
📄 xAI — ogłoszenie Grok 4.7
Terminal agentic tasks. Grok 4.7 scores 38.0% versus 20.3%.
Harvey Legal Agent Benchmark
accuracy
19.6%%
📄 xAI — ogłoszenie Grok 4.7
Benchmark of agentic legal tasks. Grok 4.7 scores 19.6% versus 15.8%.
Benchmark biobezpieczeństwa (LatchBio)
accuracy
62.4%%
📄 xAI — ogłoszenie Grok 4.7
Assessment of capability in the biosafety domain.

Pricing