OpenAI autoregressive language model (up to 1.5B parameters), decoder-only Transformer architecture, trained on the WebText dataset.
Context window
1024 tokenów
tokens
Parameters
1,5 mld (największa wersja); warianty 124M, 355M, 774M, 1,5B
parameters
Release date
14 February 2019
Access:DownloadDeployment:💻 Local☁ Cloud
Overview
Access & deployment
Download
LocalCloud
Weights: Open source
Key parameters
📏 Context: 1024 tokenów
🧩 Parameters: 1,5 mld (największa wersja); warianty 124M, 355M, 774M, 1,5B
✓ Fine-tuning
📥 Input: text
Technical specification
Context window
1024 tokenów
tokens
Parameters
1,5 mld (największa wersja); warianty 124M, 355M, 774M, 1,5B
parameters
Knowledge cutoff
1 Dec 2017
Knowledge boundary
License
Modified MIT
Hardware requirements
Variants from 124M to 1.5B parameters; the largest model (1.5B) can run on a single GPU.
Features:✓ Fine-tuning
Modalities
⬇ Input
text
⬆ Output
text
Capabilities and applications
Native model capabilities
Language modeling
Ability to predict subsequent tokens and generate coherent natural-language text based on the preceding context.
Category: language
Zero-shot learning
The model's ability to perform a new task without dataset-specific training or hyperparameter tuning — prediction is produced in a single pass from context.
Category: other
Benchmark results
6 benchmarks
LAMBADA
accuracy · zero-shot
63.24%
📄 paper
1542M model, no training or fine-tuning (Table 3).
LAMBADA
perplexity · zero-shot
8.63
📄 paper
1542M model (Table 3).
WikiText-103
perplexity · zero-shot
17.48
📄 paper
1542M model (Table 3).
WikiText-2
perplexity · zero-shot
18.34
📄 paper
1542M model (Table 3).
Penn Treebank
perplexity · zero-shot
35.76
📄 paper
1542M model (Table 3).
CBT-NE
accuracy · zero-shot
89.05%
📄 paper
1542M model (Table 3).
Technical architecture
Core Architecture
Model Form
Training Techniques
