Robots Atlas>ROBOTS ATLAS
Qwen3-4B

Qwen3-4B

Qwen3 (4B dense) · Family: Qwen3
Open (Apache 2.0) 4B-parameter LLM from Alibaba's Qwen3 series with a hybrid thinking mode and up to 128K context.
✓ Active✓ Public access⚖ Open weightsLLMReasoning model📁 Qwen3
Context window
128K (32K natywnie, do 131 072 z YaRN)
tokens
Parameters
4B (3.6B non-embedding)
parameters
Max output
38,912
tokens
Release date
29 April 2025
Access:APIDownloadHostedDeployment:💻 Local☁ Cloud📱 On-device

Overview

Qwen3-4B is an open (Apache 2.0) large language model (LLM) from the Qwen3 series developed by the Qwen team at Alibaba Cloud, released on 29 April 2025. It is a dense model with 4.0B parameters (3.6B non-embedding), built on a Transformer architecture with Grouped Query Attention (32 query heads and 8 key-value heads), SwiGLU, RoPE and RMSNorm with pre-normalization, QK-Norm stabilization and 36 layers.

A defining feature of the Qwen3 series is its hybrid operation: the model switches between a "thinking" mode for complex reasoning, math and coding, and a non-thinking mode for fast, general-purpose dialogue. The native context window is 32,768 tokens and can be extended to 131,072 tokens using the YaRN technique. The model supports over 100 languages and dialects and tool use in both modes.

Qwen3 was trained on roughly 36 trillion tokens across 119 languages, and post-training used a four-stage pipeline: long chain-of-thought cold start, reasoning-based reinforcement learning, thinking-mode fusion, and general reinforcement learning. According to the developer, even the small Qwen3-4B can rival the results of Qwen2.5-72B-Instruct. Weights are available under the Apache 2.0 license, enabling local deployment, fine-tuning and commercial use.

Classification
LLMReasoning model
Family: Qwen3
Access & deployment
APIDownloadHosted
LocalCloudOn-device
Weights: Open weights
Key parameters
📏 Context: 128K (32K natywnie, do 131 072 z YaRN)
🧩 Parameters: 4B (3.6B non-embedding)
✓ Tools · ✓ Fine-tuning
📥 Input: text

Technical specification

Context window
128K (32K natywnie, do 131 072 z YaRN)
tokens
Parameters
4B (3.6B non-embedding)
parameters
Max output tokens
38,912
tokens per response
License
Apache 2.0
Hardware requirements
About 8 GB VRAM in BF16; runs on a consumer GPU, and after quantization on CPU/edge devices.
Features:✓ Tool use✓ Fine-tuning
Modalities
⬇ Input
text
⬆ Output
textcode

Capabilities and applications

Native model capabilities
Advanced reasoning
The ability to perform multi-step, structured reasoning: analysing problems, planning steps, and drawing conclusions from hypotheses. Reasoning-first models (e.g. GPT-5.1 Thinking) dedicate a portion of inference to chains of thought before responding.
Category: reasoning
Extended thinking mode
A reasoning-model variant with a larger inference budget: more thinking cycles, higher answer precision at the cost of response time. Choice between 'standard' and 'extended' thinking is left to the user (e.g. the selector in GPT-5.2 Pro).
Category: reasoning
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Mathematical reasoning
The model's ability to solve mathematical tasks requiring multi-step reasoning — equations, proofs, combinatorics, geometry, calculus and competition-level problems.
Category: reasoning
Instruction following
Precisely following instructions contained in the prompt: response format, length, style, constraints (e.g. 'reply in six words'). GPT-5.1 significantly improved this capability compared to GPT-5.
Category: language
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Function Calling
Category: planning

Benchmark results

8 benchmarks
AIME 2024
accuracy · thinking mode
73.8%
📄 Qwen3 Technical Report (arXiv:2505.09388)
AIME 2025
accuracy · thinking mode
65.6%
📄 Qwen3 Technical Report (arXiv:2505.09388)
LiveCodeBench
pass rate · thinking mode
54.2%
📄 Qwen3 Technical Report (arXiv:2505.09388)
GPQA
accuracy · thinking mode
55.9%
📄 Qwen3 Technical Report (arXiv:2505.09388)
MMLU-Redux
accuracy · thinking mode
83.7%
📄 Qwen3 Technical Report (arXiv:2505.09388)
MMLU-Pro
accuracy · thinking mode
50.58%
📄 Qwen3 Technical Report (arXiv:2505.09388)
Arena-Hard
win rate · thinking mode
76.6%
📄 Qwen3 Technical Report (arXiv:2505.09388)
BBH (BIG-Bench Hard)
accuracy · thinking mode
72.59%
📄 Qwen3 Technical Report (arXiv:2505.09388)

Technical architecture