The smallest dense model in Alibaba's Qwen3 family (0.6B parameters) with a hybrid thinking mode and open weights under the Apache 2.0 license.
Context window
32K tokenów (natywnie 32 768)
tokens
Parameters
0.6B (0.44B bez embeddingów)
parameters
Release date
29 April 2025
Access:DownloadAPIHostedDeployment:💻 Local☁ Cloud📱 On-device
Overview
Access & deployment
DownloadAPIHosted
LocalCloudOn-device
Weights: Open source
Key parameters
📏 Context: 32K tokenów (natywnie 32 768)
🧩 Parameters: 0.6B (0.44B bez embeddingów)
✓ Tools · ✓ Fine-tuning
📥 Input: text
Technical specification
Context window
32K tokenów (natywnie 32 768)
tokens
Parameters
0.6B (0.44B bez embeddingów)
parameters
License
Apache 2.0
Hardware requirements
Very low — 0.6B parameters fit in a few hundred MB (quantized) and run on CPU, mobile devices and GPUs with 2–4 GB VRAM.
Features:✓ Tool use✓ Fine-tuning
Modalities
⬇ Input
text
⬆ Output
textcodestructured_data
Capabilities and applications
Native model capabilities
Language modeling
Ability to predict subsequent tokens and generate coherent natural-language text based on the preceding context.
Category: language
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Instruction following
Precisely following instructions contained in the prompt: response format, length, style, constraints (e.g. 'reply in six words'). GPT-5.1 significantly improved this capability compared to GPT-5.
Category: language
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Function Calling
Category: planning
Structured output
Producing data in structured formats such as JSON.
Category: structured_generation
Extended thinking mode
A reasoning-model variant with a larger inference budget: more thinking cycles, higher answer precision at the cost of response time. Choice between 'standard' and 'extended' thinking is left to the user (e.g. the selector in GPT-5.2 Pro).
Category: reasoning
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Technical architecture
Core Architecture
Model Form
Training Techniques
