Robots Atlas>ROBOTS ATLAS
GPT-4o

GPT-4o

4o · Family: GPT
OpenAI's natively multimodal model (text, image, audio), released May 13, 2024. Real-time voice (~320 ms), 128K context, function calling, 50+ languages.
✓ Active✓ Public accessMultimodalLLMTool-using modelVisionAudioAudio📁 GPT
Context window
128K
tokens
Parameters
Nieujawnione
parameters
Max output
16,384
tokens
Release date
13 May 2024
Access:APIHostedDeployment:☁ Cloud

Overview

Overview

GPT-4o („o” for omni) is OpenAI’s natively multimodal model unveiled on May 13, 2024. A single model handles understanding and generation across text, image, and audio, enabling real-time spoken conversation with an average latency of about 320 ms — close to human response time.

Compared with GPT-4 Turbo it delivers flagship-level quality at a substantially lower price and higher throughput. It supports a 128K-token context window, function calling, Structured Outputs, and more than 50 languages. The model’s knowledge cutoff is October 2023.

Use cases

It excels at voice and text assistants, coding, document and image analysis (OCR, charts), live translation, summarization, and agentic tool-use tasks.

Classification
MultimodalLLMTool-using modelVisionAudioAudio
Family: GPT
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📏 Context: 128K
🧩 Parameters: Nieujawnione
Tools · ✓ Fine-tuning
📥 Input: text, image, audio

Technical specification

Context window
128K
tokens
Parameters
Nieujawnione
parameters
Max output tokens
16,384
tokens per response
Knowledge cutoff
1 Oct 2023
Knowledge boundary
License
Proprietary (OpenAI)
Hardware requirements
Closed model, available only via the OpenAI API and Azure OpenAI (cloud). No weights available to self-host locally.
Features:Tool useFine-tuning
Modalities
⬇ Input
textimageaudio
⬆ Output
textimageaudiocode

Capabilities and applications

Native model capabilities
Multimodal understanding
Category: multimodal
Image understanding
Analysing and interpreting the content of images.
Category: vision
Audio understanding
Category: audio
Video Understanding
Category: video
Voice Conversation
Ability to conduct multi-turn real-time voice conversations with context retention and natural speech pacing.
Category: speech
Speech to text
Category: speech
Text to speech
Category: speech
Real-time inference
The model's ability to generate responses with very low latency (>1000 tokens/sec) on specialized inference hardware (e.g. Cerebras WSE), enabling interactive, turn-by-turn collaboration with a human.
Category: coding
Natural conversation
Conducting a conversation with a tone close to human: a warmer voice, empathy in emotional responses, humour, and avoiding the stiff 'AI assistant jargon'. Introduced as a deliberate improvement in GPT-5.1 Instant.
Category: language
Live Translation
Real-time speech translation between multiple languages without interrupting the audio stream.
Category: speech
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Mathematical reasoning
The model's ability to solve mathematical tasks requiring multi-step reasoning — equations, proofs, combinatorics, geometry, calculus and competition-level problems.
Category: reasoning
OCR
Recognising text within images and documents.
Category: vision
Chart understanding
Reading and interpreting charts, tables and diagrams.
Category: vision
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Instruction following
Precisely following instructions contained in the prompt: response format, length, style, constraints (e.g. 'reply in six words'). GPT-5.1 significantly improved this capability compared to GPT-5.
Category: language
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Function Calling
Category: planning
Parallel Tool Calls
Ability to invoke multiple external tools simultaneously while generating a response.
Category: reasoning
Structured output
Producing data in structured formats such as JSON.
Category: structured_generation
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Streaming output
Category: reasoning
Prompt caching
Cost-performance optimisation: repeated prompt fragments (e.g. system prompt, long documentation) are cached server-side and cheaper in subsequent calls. Significantly reduces cost for applications with long contexts.
Category: other
Interleaved Multimodal Input
Category: reasoning

Benchmark results

7 benchmarks
MMLU
accuracy · 0-shot CoT
88.7%
📅 6 Aug 2024📄 openai/simple-evals
Snapshot gpt-4o-2024-08-06.
GPQA
accuracy · GPQA Diamond, 0-shot CoT
53.1%
📅 6 Aug 2024📄 openai/simple-evals
Snapshot gpt-4o-2024-08-06.
MATH
accuracy · 0-shot CoT
75.9%
📅 6 Aug 2024📄 openai/simple-evals
Snapshot gpt-4o-2024-08-06.
HumanEval
pass@1 · 0-shot
90.2%
📅 6 Aug 2024📄 openai/simple-evals
Snapshot gpt-4o-2024-08-06.
MGSM
accuracy · 0-shot CoT
90.0%
📅 6 Aug 2024📄 openai/simple-evals
Snapshot gpt-4o-2024-08-06.
DROP
F1 · 3-shot
79.8%
📅 6 Aug 2024📄 openai/simple-evals
Snapshot gpt-4o-2024-08-06.
SimpleQA
accuracy · 0-shot
40.1%
📅 6 Aug 2024📄 openai/simple-evals
Snapshot gpt-4o-2024-08-06.

Pricing

Technical architecture

Deployment and security

🔒 Security / Enterprise
✓ Verified enterprise information
Updated: 22 Jul 2026↗ Security documentation