Alibaba Qwen3-14B language model, 14.8B parameters, Apache 2.0. Hybrid thinking/non-thinking mode, 128K context, 119 languages, strong in math, coding and agent tasks.
Context window
128K
tokens
Parameters
14.8B
parameters
Max output
32,768
tokens
Release date
29 April 2025
Access:DownloadAPIHostedDeployment:๐ป Localโ Cloud๐ฑ On-device
Overview
Access & deployment
DownloadAPIHosted
LocalCloudOn-device
Weights: Open source
Key parameters
๐ Context: 128K
๐งฉ Parameters: 14.8B
โ Toolsย ยทย โ Fine-tuning
๐ฅ Input: text
Technical specification
Context window
128K
tokens
Parameters
14.8B
parameters
Max output tokens
32,768
tokens per response
License
Apache 2.0
Hardware requirements
GPU with ~32 GB VRAM for BF16 (14.8B params, 40 layers), or ~12 GB with Q4 quantization. Flash Attention 2 recommended. Supported: Transformers (>=4.51.0), vLLM (>=0.8.5), SGLang (>=0.4.6.post1), llama.cpp, Ollama, LMStudio.
Features:โ Tool useโ Fine-tuning
Modalities
โฌ Input
text
โฌ Output
textcode
Capabilities and applications
Native model capabilities
Reasoning
The model's ability to perform multi-step logical inference, solve complex problems and decompose tasks into steps.
Category: reasoning
Coding
Generating, completing, explaining and debugging code across multiple programming languages.
Category: coding
Multilingual
Understanding and generating text in many languages and translating between them.
Category: language
Long context
Processing very long inputs (tens to hundreds of thousands of tokens) while maintaining coherence.
Category: language
Multi-step reasoning
Carrying out multi-step chains of reasoning across long, complex tasks.
Category: reasoning
Mathematical reasoning
The model's ability to solve mathematical tasks requiring multi-step reasoning โ equations, proofs, combinatorics, geometry, calculus and competition-level problems.
Category: reasoning
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Planning
Forming and executing action plans for complex tasks.
Category: planning
Function Calling
Category: planning
Structured output
Producing data in structured formats such as JSON.
Category: structured_generation
Language modeling
Ability to predict subsequent tokens and generate coherent natural-language text based on the preceding context.
Category: language
Application domains
Benchmark results
10 benchmarks
MMLU
accuracy ยท 5-shot, base model Qwen3-14B-Base
81.05%
๐
14 May 2025๐ Qwen3 Technical Report, Table 5 (arXiv 2505.09388)
MMLU-Pro
accuracy ยท 5-shot CoT, base model Qwen3-14B-Base
61.03%
๐
14 May 2025๐ Qwen3 Technical Report, Table 5 (arXiv 2505.09388)
GPQA
accuracy ยท base model Qwen3-14B-Base
39.90%
๐
14 May 2025๐ Qwen3 Technical Report, Table 5 (arXiv 2505.09388)
MATH
accuracy ยท 4-shot CoT, base model Qwen3-14B-Base
62.02%
๐
14 May 2025๐ Qwen3 Technical Report, Table 5 (arXiv 2505.09388)
GSM8K
accuracy ยท 4-shot CoT, base model Qwen3-14B-Base
92.49%
๐
14 May 2025๐ Qwen3 Technical Report, Table 5 (arXiv 2505.09388)
MGSM
accuracy ยท 8-shot CoT, multilingual math, base model Qwen3-14B-Base
79.20%
๐
14 May 2025๐ Qwen3 Technical Report, Table 5 (arXiv 2505.09388)
BBH (BIG-Bench Hard)
accuracy ยท 3-shot, base model Qwen3-14B-Base
81.07%
๐
14 May 2025๐ Qwen3 Technical Report, Table 5 (arXiv 2505.09388)
MBPP
pass@1 ยท 3-shot, coding, base model Qwen3-14B-Base
73.40%
๐
14 May 2025๐ Qwen3 Technical Report, Table 5 (arXiv 2505.09388)
EvalPlus
pass@1 ยท coding, base model Qwen3-14B-Base
72.23%
๐
14 May 2025๐ Qwen3 Technical Report, Table 5 (arXiv 2505.09388)
MultiPL-E
pass@1 ยท multilingual coding, base model Qwen3-14B-Base
61.69%
๐
14 May 2025๐ Qwen3 Technical Report, Table 5 (arXiv 2505.09388)
Pricing
Technical architecture
Model Form
Training Techniques
