Anthropic's Claude-family large language model (5.5, September 2026) for long-running agentic coding, computer use and knowledge work. 1M-token context window.
Context window
1M
tokens
Max output
128,000
tokens
Release date
22 September 2026
Access:APIHostedDeployment:☁ Cloud
Overview
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📏 Context: 1M
✓ Tools
📥 Input: text, image
Technical specification
Context window
1M
tokens
Max output tokens
128,000
tokens per response
Knowledge cutoff
1 Jun 2026
Knowledge boundary
License
Proprietary (Anthropic Commercial Terms of Service)
Hardware requirements
Closed model, available only via cloud (Claude API, Amazon Bedrock, Google Cloud Vertex AI, Microsoft Foundry). No local deployment.
Features:✓ Tool use
Modalities
⬇ Input
textimage
⬆ Output
textcode
Capabilities and applications
Native model capabilities
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Advanced reasoning
The ability to perform multi-step, structured reasoning: analysing problems, planning steps, and drawing conclusions from hypotheses. Reasoning-first models (e.g. GPT-5.1 Thinking) dedicate a portion of inference to chains of thought before responding.
Category: reasoning
Multi-step reasoning
Carrying out multi-step chains of reasoning across long, complex tasks.
Category: reasoning
Extended thinking mode
A reasoning-model variant with a larger inference budget: more thinking cycles, higher answer precision at the cost of response time. Choice between 'standard' and 'extended' thinking is left to the user (e.g. the selector in GPT-5.2 Pro).
Category: reasoning
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Agentic coding
Multi-hour, multi-step programming tasks performed autonomously by the model: cloning a repository, running tests, iterating on fixes, integrating with CLI tools. Characteristic of Codex variants (GPT-5.1-Codex-Mini, Codex-Max).
Category: coding
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Computer use
The model's ability to operate a computer interface by interpreting screenshots and generating actions such as clicks, typing, and navigating applications.
Category: planning
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Structured output
Producing data in structured formats such as JSON.
Category: structured_generation
Image understanding
Analysing and interpreting the content of images.
Category: vision
Chart understanding
Reading and interpreting charts, tables and diagrams.
Category: vision
OCR
Recognising text within images and documents.
Category: vision
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Benchmark results
8 benchmarks
Terminal-Bench 4.0
accuracy
66.4%
📄 Anthropic – oficjalna zapowiedź (anthropic.com/claude/opus)
FrontierCode v1.1
accuracy
54.4%
📄 Anthropic – oficjalna zapowiedź (anthropic.com/claude/opus)
CursorBench 4.0
accuracy
57.8%
📄 Anthropic – oficjalna zapowiedź (anthropic.com/claude/opus)
GDPval-AA v2.1
Elo
1846Elo
📄 Anthropic – oficjalna zapowiedź (anthropic.com/claude/opus)
AutomationBench
accuracy
40%
📄 Anthropic – oficjalna zapowiedź (anthropic.com/claude/opus)
Humanity's Last Exam
accuracy · with tools
67.7%
📄 Anthropic – oficjalna zapowiedź (anthropic.com/claude/opus)
OSWorld 2.0
accuracy
81.8%
📄 Anthropic – oficjalna zapowiedź (anthropic.com/claude/opus)
Wynik częściowy (partial), wg oficjalnej zapowiedzi.
Chartography
accuracy · with tools
89%
📄 Anthropic – oficjalna zapowiedź (anthropic.com/claude/opus)
Pricing
Technical architecture
Core Architecture
Training Techniques
Deployment and security
🔒 Security / Enterprise
✓ Verified enterprise information
