
OpenAI o-series reasoning model trained with reinforcement learning; builds an internal chain of thought before answering. Excels at math, science and coding.
Context window
200K
tokens
Max output
100,000
tokens
Release date
5 December 2024
Access:APIHostedDeployment:☁ Cloud
Overview
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📏 Context: 200K
✓ Tools
📥 Input: text, image
Platforms
Technical specification
Context window
200K
tokens
Max output tokens
100,000
tokens per response
Knowledge cutoff
1 Oct 2023
Knowledge boundary
License
Proprietary
Features:✓ Tool use
Modalities
⬇ Input
textimage
⬆ Output
text
Capabilities and applications
Native model capabilities
Advanced reasoning
The ability to perform multi-step, structured reasoning: analysing problems, planning steps, and drawing conclusions from hypotheses. Reasoning-first models (e.g. GPT-5.1 Thinking) dedicate a portion of inference to chains of thought before responding.
Category: reasoning
Mathematical reasoning
The model's ability to solve mathematical tasks requiring multi-step reasoning — equations, proofs, combinatorics, geometry, calculus and competition-level problems.
Category: reasoning
Multi-step reasoning
Carrying out multi-step chains of reasoning across long, complex tasks.
Category: reasoning
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Image understanding
Analysing and interpreting the content of images.
Category: vision
Structured output
Producing data in structured formats such as JSON.
Category: structured_generation
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Application domains
Benchmark results
3 benchmarks
AIME 2024
accuracy · single sample
83%
📅 12 Sept 2024📄 OpenAI (Learning to Reason with LLMs)
Codeforces
percentile
89th percentile
📅 12 Sept 2024📄 OpenAI (Learning to Reason with LLMs)
GPQA
accuracy
📄 OpenAI
On GPQA Diamond, o1 exceeds the accuracy of PhD-level human experts in physics, chemistry and biology.
Pricing
Technical architecture
Core Architecture
Model Form
Deployment and security
☁ Available on platforms
🔒 Security / Enterprise
✓ Verified enterprise information
Updated: 23 Jul 2026↗ Security documentation
Sources and related pages
4 sources