An open (Apache 2.0) MoE model from the Qwen3 family (Alibaba): 30.5B total parameters, 3.3B active (128 experts, 8 activated); thinking/non-thinking modes, context up to 262K, strong at reasoning, coding and agentic tasks.
Context window
262K (warianty 2507, natywnie); bazowy A3B: 32K natywnie / 131K z YaRN
tokens
Parameters
30.5B total / 3.3B active (MoE: 128 ekspertow, 8 aktywnych)
parameters
Max output
32,768
tokens
Release date
29 April 2025
Access:DownloadAPIHostedDeployment:๐ป Localโ Cloud๐ฑ On-device
Overview
Applications
Access & deployment
DownloadAPIHosted
LocalCloudOn-device
Weights: Open source
Key parameters
๐ Context: 262K (warianty 2507, natywnie); bazowy A3B: 32K natywnie / 131K z YaRN
๐งฉ Parameters: 30.5B total / 3.3B active (MoE: 128 ekspertow, 8 aktywnych)
โ Toolsย ยทย โ Fine-tuning
๐ฅ Input: text
Platforms
Technical specification
Context window
262K (warianty 2507, natywnie); bazowy A3B: 32K natywnie / 131K z YaRN
tokens
Parameters
30.5B total / 3.3B active (MoE: 128 ekspertow, 8 aktywnych)
parameters
Max output tokens
32,768
tokens per response
Knowledge cutoff
1 Apr 2025
Knowledge boundary
License
Apache 2.0
Hardware requirements
MoE with 3.3B active - efficient at inference; BF16 weights ~61 GB (multi-GPU self-host or quantization). Supported: Transformers (>=4.51.0), vLLM (>=0.8.5), SGLang (>=0.4.6.post1), llama.cpp, Ollama, LMStudio, MLX.
Features:โ Tool useโ Fine-tuning
Modalities
โฌ Input
text
โฌ Output
textcode
Capabilities and applications
Native model capabilities
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Multi-step reasoning
Carrying out multi-step chains of reasoning across long, complex tasks.
Category: reasoning
Mathematical reasoning
The model's ability to solve mathematical tasks requiring multi-step reasoning โ equations, proofs, combinatorics, geometry, calculus and competition-level problems.
Category: reasoning
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Long context
Support for large context windows โ tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Planning
Forming and executing action plans for complex tasks.
Category: planning
Function Calling
Category: planning
Structured output
Producing data in structured formats such as JSON.
Category: structured_generation
Language modeling
Ability to predict subsequent tokens and generate coherent natural-language text based on the preceding context.
Category: language
Pricing
Technical architecture
Core Architecture
Model Form
Training Techniques
Related Technologies
Deployment and security
โ Available on platforms
๐ Security / Enterprise
โ Verified enterprise information
Open-weight model under Apache 2.0 - self-hosting / on-premises possible with full data control. Available on Hugging Face; supports function calling and tool integration.
Updated: 31 Jul 2026โ Security documentation
