Instruction-tuned 1.5B-parameter language model from Alibaba's Qwen2 series. Transformer with GQA and SwiGLU, 32K context window; weights under Apache 2.0.
Context window
32K
tokens
Parameters
1.5B
parameters
Release date
7 June 2024
Access:DownloadAPIDeployment:💻 Local☁ Cloud📱 On-device
Overview
Access & deployment
DownloadAPI
LocalCloudOn-device
Weights: Open source
Key parameters
📏 Context: 32K
🧩 Parameters: 1.5B
✓ Tools · ✓ Fine-tuning
📥 Input: text
Technical specification
Context window
32K
tokens
Parameters
1.5B
parameters
License
Apache 2.0
Hardware requirements
About 1.5B parameters; in BF16 it needs roughly 3–4 GB of memory. Runs locally on a single consumer GPU or CPU; quantized (INT4/GGUF) it runs on edge devices.
Features:✓ Tool use✓ Fine-tuning
Modalities
⬇ Input
text
⬆ Output
textcode
Capabilities and applications
Native model capabilities
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Mathematical reasoning
The model's ability to solve mathematical tasks requiring multi-step reasoning — equations, proofs, combinatorics, geometry, calculus and competition-level problems.
Category: reasoning
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Instruction following
Precisely following instructions contained in the prompt: response format, length, style, constraints (e.g. 'reply in six words'). GPT-5.1 significantly improved this capability compared to GPT-5.
Category: language
Function Calling
Category: planning
Benchmark results
5 benchmarks
MMLU
accuracy
52.4%
📅 7 Jun 2024📄 documentation
HumanEval
pass@1
37.8%
📅 7 Jun 2024📄 documentation
GSM8K
accuracy
61.6%
📅 7 Jun 2024📄 documentation
C-Eval
accuracy
63.8%
📅 7 Jun 2024📄 documentation
IFEval
strict prompt accuracy
29.0%
📅 7 Jun 2024📄 documentation
Technical architecture
Core Architecture
Model Form
Training Techniques
