OpenAI's 175-billion-parameter language model (2020) that popularized few-shot in-context learning without weight fine-tuning.
Context window
2048 tokenów
tokens
Parameters
175B
parameters
Max output
2,048
tokens
Release date
28 May 2020
Access:APIHostedDeployment:☁ Cloud
Overview
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📏 Context: 2048 tokenów
🧩 Parameters: 175B
✓ Fine-tuning
📥 Input: text
Technical specification
Context window
2048 tokenów
tokens
Parameters
175B
parameters
Max output tokens
2,048
tokens per response
License
Proprietary (OpenAI API)
Hardware requirements
Training: ~3.14×10^23 FLOPs on a Microsoft Azure V100 GPU cluster. The 16-bit model needs ~350 GB. Weights are closed — inference only via the OpenAI API (cloud).
Features:✓ Fine-tuning
Modalities
⬇ Input
text
⬆ Output
text
Capabilities and applications
Native model capabilities
Language modeling
Ability to predict subsequent tokens and generate coherent natural-language text based on the preceding context.
Category: language
Zero-shot learning
The model's ability to perform a new task without dataset-specific training or hyperparameter tuning — prediction is produced in a single pass from context.
Category: other
Few-shot learning
The model's ability to perform a new task from a handful of examples provided directly in the prompt, without any weight updates or fine-tuning.
Category: other
Application domains
Benchmark results
6 benchmarks
LAMBADA
accuracy · few-shot (64-shot)
86.4%
📅 28 May 2020📄 paper
Score reported in "Language Models are Few-Shot Learners" (GPT-3 175B).
HellaSwag
accuracy · few-shot
79.3%
📅 28 May 2020📄 paper
Score reported in "Language Models are Few-Shot Learners" (GPT-3 175B).
TriviaQA
accuracy · few-shot
71.2%
📅 28 May 2020📄 paper
Score reported in "Language Models are Few-Shot Learners" (GPT-3 175B).
WinoGrande
accuracy · few-shot
77.7%
📅 28 May 2020📄 paper
Score reported in "Language Models are Few-Shot Learners" (GPT-3 175B).
PIQA
accuracy · few-shot
82.8%
📅 28 May 2020📄 paper
Score reported in "Language Models are Few-Shot Learners" (GPT-3 175B).
ARC-Challenge
accuracy · few-shot
51.5%
📅 28 May 2020📄 paper
Score reported in "Language Models are Few-Shot Learners" (GPT-3 175B).
Pricing
Technical architecture
Core Architecture
Model Form
Training Techniques
