Robots Atlas>ROBOTS ATLAS
GPT-4

GPT-4

4 · Family: GPT
Multimodal large language model by OpenAI, released March 2023 as successor to GPT-3.5. Accepts text and images, outputs text. Closed weights; via API and formerly ChatGPT.
⚠ Deprecated✓ Public accessLLMMultimodal📁 GPT
Context window
8K (wariant 32K)
tokens
Parameters
Nieujawnione
parameters
Release date
14 March 2023
Access:APIHostedDeployment:☁ Cloud

Overview

GPT-4 is a multimodal large language model (LLM) developed by OpenAI and released on 14 March 2023 as the successor to GPT-3.5. It accepts text and image inputs and produces text outputs.

GPT-4 was offered with an 8,192-token context window and an extended 32,768-token variant. OpenAI did not disclose the parameter count or architecture details; the model is based on the Transformer architecture and was tuned using reinforcement learning from human feedback (RLHF).

The model achieved strong results on professional and academic exams — including roughly the 90th percentile on a simulated Uniform Bar Exam — and high scores on the MMLU, HumanEval and GSM8K benchmarks. Model weights remain closed; access is provided through the OpenAI API and Azure OpenAI, and historically through ChatGPT (Plus).

Classification
LLMMultimodal
Family: GPT
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📏 Context: 8K (wariant 32K)
🧩 Parameters: Nieujawnione
Tools
📥 Input: text, image

Technical specification

Context window
8K (wariant 32K)
tokens
Parameters
Nieujawnione
parameters
Knowledge cutoff
1 Sept 2021
Knowledge boundary
License
Proprietary (zamknięta)
Hardware requirements
Closed model, available only as a cloud service (OpenAI API, Azure OpenAI). No local deployment.
Features:Tool use
Modalities
⬇ Input
textimage
⬆ Output
textcode

Capabilities and applications

Native model capabilities
Language modeling
Ability to predict subsequent tokens and generate coherent natural-language text based on the preceding context.
Category: language
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Image understanding
Analysing and interpreting the content of images.
Category: vision
Multimodal understanding
Category: multimodal
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Instruction following
Precisely following instructions contained in the prompt: response format, length, style, constraints (e.g. 'reply in six words'). GPT-5.1 significantly improved this capability compared to GPT-5.
Category: language
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Mathematical reasoning
The model's ability to solve mathematical tasks requiring multi-step reasoning — equations, proofs, combinatorics, geometry, calculus and competition-level problems.
Category: reasoning
Function Calling
Category: planning

Benchmark results

5 benchmarks
MMLU
accuracy · 5-shot
86.4%
📄 GPT-4 Technical Report (OpenAI, 2023)
HumanEval
pass@1 · zero-shot
67.0%
📄 GPT-4 Technical Report (OpenAI, 2023)
GSM8K
accuracy · 5-shot, chain-of-thought
92.0%
📄 GPT-4 Technical Report (OpenAI, 2023)
HellaSwag
accuracy · 10-shot
95.3%
📄 GPT-4 Technical Report (OpenAI, 2023)
DROP
f1 · 3-shot
80.9
📄 GPT-4 Technical Report (OpenAI, 2023)

Pricing

Technical architecture

Deployment and security

🔒 Security / Enterprise
✓ Verified enterprise information
Updated: 10 Sept 2026↗ Security documentation