Robots Atlas>ROBOTS ATLAS
Qwen3-30B-A3B

Qwen3-30B-A3B

30B-A3Bย ยทย Family: Qwen3
An open (Apache 2.0) MoE model from the Qwen3 family (Alibaba): 30.5B total parameters, 3.3B active (128 experts, 8 activated); thinking/non-thinking modes, context up to 262K, strong at reasoning, coding and agentic tasks.
โœ“ Activeโœ“ Public accessโš– Open sourceLLMReasoning modelTool-using model๐Ÿ“ Qwen3
Context window
262K (warianty 2507, natywnie); bazowy A3B: 32K natywnie / 131K z YaRN
tokens
Parameters
30.5B total / 3.3B active (MoE: 128 ekspertow, 8 aktywnych)
parameters
Max output
32,768
tokens
Release date
29 April 2025
Access:DownloadAPIHostedDeployment:๐Ÿ’ป Localโ˜ Cloud๐Ÿ“ฑ On-device

Overview

Qwen3-30B-A3B is an open language model from the Qwen3 family developed by the Qwen team (Alibaba), built on a Mixture-of-Experts (MoE) architecture. It has 30.5B total parameters, of which only 3.3B are activated per token (128 experts, 8 activated), delivering large-model quality at the compute cost of a much smaller one.

The model supports two modes: thinking (explicit step-by-step reasoning for complex tasks - math, code, logic) and non-thinking (fast general-purpose responses), switchable within a single model. It stands out for strong reasoning, coding, agentic capabilities (function calling, tool use) and multilinguality.

The base Qwen3-30B-A3B supports 32K context natively and 131K with YaRN; the newer 2507 variants (Instruct and Thinking) support 262,144 tokens natively. Weights are released under Apache 2.0 (self-hosting), and the model is available via Hugging Face and provider APIs (e.g. OpenRouter, Alibaba DashScope). Supported by Transformers, vLLM, SGLang, llama.cpp and Ollama.

Classification
LLMReasoning modelTool-using model
Family: Qwen3
Access & deployment
DownloadAPIHosted
LocalCloudOn-device
Weights: Open source
Key parameters
๐Ÿ“ Context: 262K (warianty 2507, natywnie); bazowy A3B: 32K natywnie / 131K z YaRN
๐Ÿงฉ Parameters: 30.5B total / 3.3B active (MoE: 128 ekspertow, 8 aktywnych)
โœ“ Toolsย ยทย โœ“ Fine-tuning
๐Ÿ“ฅ Input: text

Technical specification

Context window
262K (warianty 2507, natywnie); bazowy A3B: 32K natywnie / 131K z YaRN
tokens
Parameters
30.5B total / 3.3B active (MoE: 128 ekspertow, 8 aktywnych)
parameters
Max output tokens
32,768
tokens per response
Knowledge cutoff
1 Apr 2025
Knowledge boundary
License
Apache 2.0
Hardware requirements
MoE with 3.3B active - efficient at inference; BF16 weights ~61 GB (multi-GPU self-host or quantization). Supported: Transformers (>=4.51.0), vLLM (>=0.8.5), SGLang (>=0.4.6.post1), llama.cpp, Ollama, LMStudio, MLX.
Features:โœ“ Tool useโœ“ Fine-tuning
Modalities
โฌ‡ Input
text
โฌ† Output
textcode

Capabilities and applications

Native model capabilities
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Multi-step reasoning
Carrying out multi-step chains of reasoning across long, complex tasks.
Category: reasoning
Mathematical reasoning
The model's ability to solve mathematical tasks requiring multi-step reasoning โ€” equations, proofs, combinatorics, geometry, calculus and competition-level problems.
Category: reasoning
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Long context
Support for large context windows โ€” tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Planning
Forming and executing action plans for complex tasks.
Category: planning
Function Calling
Category: planning
Structured output
Producing data in structured formats such as JSON.
Category: structured_generation
Language modeling
Ability to predict subsequent tokens and generate coherent natural-language text based on the preceding context.
Category: language

Pricing

Technical architecture

Deployment and security

๐Ÿ”’ Security / Enterprise
โœ“ Verified enterprise information

Open-weight model under Apache 2.0 - self-hosting / on-premises possible with full data control. Available on Hugging Face; supports function calling and tool integration.

Updated: 31 Jul 2026โ†— Security documentation