Anthropic's Claude family AI model focused on coding, agentic tasks and computer use, with a 200K-token context window.
Context window
200K
tokens
Max output
64,000
tokens
Release date
29 September 2025
Access:APIHostedDeployment:โ Cloud
Overview
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
๐ Context: 200K
โ Tools
๐ฅ Input: text, image
Technical specification
Context window
200K
tokens
Max output tokens
64,000
tokens per response
Features:โ Tool use
Modalities
โฌ Input
textimage
โฌ Output
textcode
Capabilities and applications
Native model capabilities
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Multi-step reasoning
Carrying out multi-step chains of reasoning across long, complex tasks.
Category: reasoning
Long context
Support for large context windows โ tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Image understanding
Analysing and interpreting the content of images.
Category: vision
Function Calling
Category: planning
Parallel Tool Calls
Ability to invoke multiple external tools simultaneously while generating a response.
Category: reasoning
Planning
Forming and executing action plans for complex tasks.
Category: planning
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Computer use
The model's ability to operate a computer interface by interpreting screenshots and generating actions such as clicks, typing, and navigating applications.
Category: planning
Benchmark results
2 benchmarks
SWE-bench
accuracy ยท SWE-bench Verified, 200K thinking budget, averaged over 10 trials
77.2%
๐
29 Sept 2025๐ Anthropic announcement (Introducing Claude Sonnet 4.5)
OSWorld
accuracy ยท OSWorld-Verified framework, 100 max steps, averaged across 4 runs
61.4%
๐
29 Sept 2025๐ Anthropic announcement (Introducing Claude Sonnet 4.5)
Pricing
Technical architecture
Core Architecture
Training Techniques
