Xiaomi's series of omnimodal agentic models (Pro-RL, Flash-RL, Distill-Qwen-9B) trained with scaled RL; text/image/video/audio input, 1M context.
Context window
1M tokenów
tokens
Parameters
1,02 bln / 42 mld aktywnych (Pro-RL); 309 mld / 15 mld (Flash-RL); 9 mld (Distill-Qwen-9B)
parameters
Release date
21 September 2026
Access:APIDownloadHostedDeployment:☁ Cloud💻 Local
Overview
Applications
Access & deployment
APIDownloadHosted
CloudLocal
Weights: Open source
Key parameters
📏 Context: 1M tokenów
🧩 Parameters: 1,02 bln / 42 mld aktywnych (Pro-RL); 309 mld / 15 mld (Flash-RL); 9 mld (Distill-Qwen-9B)
✓ Tools · ✓ Fine-tuning
📥 Input: text, image, video, audio
Technical specification
Context window
1M tokenów
tokens
Parameters
1,02 bln / 42 mld aktywnych (Pro-RL); 309 mld / 15 mld (Flash-RL); 9 mld (Distill-Qwen-9B)
parameters
License
MIT
Hardware requirements
Large variants (Pro-RL, Flash-RL) require multi-GPU deployment with tensor parallelism; SGLang or vLLM recommended. The 9B Distill-Qwen-9B variant can run on a single GPU.
Features:✓ Tool use✓ Fine-tuning
Modalities
⬇ Input
textimagevideoaudio
⬆ Output
textcode
Capabilities and applications
Native model capabilities
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Agentic coding
Multi-hour, multi-step programming tasks performed autonomously by the model: cloning a repository, running tests, iterating on fixes, integrating with CLI tools. Characteristic of Codex variants (GPT-5.1-Codex-Mini, Codex-Max).
Category: coding
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Multimodal understanding
Category: multimodal
Image understanding
Analysing and interpreting the content of images.
Category: vision
Video understanding
The model's ability to analyse and interpret video content — recognising actions, motion, events and relationships between objects over time.
Category: video
Audio understanding
Category: audio
Cybersecurity
The model ability to perform computer-security tasks: vulnerability analysis, proof-of-concept exploit generation, patching, and cybersecurity question answering.
Category: other
Computer use
The model's ability to operate a computer interface by interpreting screenshots and generating actions such as clicks, typing, and navigating applications.
Category: planning
Benchmark results
7 benchmarks
DeepSWE v1.1
Code Agent (MiMo-V2.6 Pro)
71.9
📅 21 Sept 2026📄 Raport techniczny MiMo-V2.6
Terminal Bench 2.1
General Agent (MiMo-V2.6 Pro)
89.9
📅 21 Sept 2026📄 Raport techniczny MiMo-V2.6
Toolathlon-Verified
General Agent (MiMo-V2.6 Pro)
76.9
📅 21 Sept 2026📄 Raport techniczny MiMo-V2.6
OSWorld
General Agent (MiMo-V2.6 Pro)
82.0
📅 21 Sept 2026📄 Raport techniczny MiMo-V2.6
GDPval-AA
General Agent (MiMo-V2.6 Pro)
1673
📅 21 Sept 2026📄 Raport techniczny MiMo-V2.6
Agents' Last Exam
General Agent (MiMo-V2.6 Pro)
31.6
📅 21 Sept 2026📄 Raport techniczny MiMo-V2.6
CyberGym
Cybersecurity (MiMo-V2.6 Pro)
94.0
📅 21 Sept 2026📄 Raport techniczny MiMo-V2.6
