Robots Atlas>ROBOTS ATLAS
MiMo-V2.6

MiMo-V2.6

V2.6 · Family: MiMo
Xiaomi's series of omnimodal agentic models (Pro-RL, Flash-RL, Distill-Qwen-9B) trained with scaled RL; text/image/video/audio input, 1M context.
✓ Active✓ Public access⚖ Open sourceMultimodalReasoning modelTool-using model📁 MiMo
Context window
1M tokenów
tokens
Parameters
1,02 bln / 42 mld aktywnych (Pro-RL); 309 mld / 15 mld (Flash-RL); 9 mld (Distill-Qwen-9B)
parameters
Release date
21 September 2026
Access:APIDownloadHostedDeployment:☁ Cloud💻 Local

Overview

MiMo-V2.6 is a series of natively omnimodal language models developed by the Xiaomi MiMo team and released on 21 September 2026 under the MIT license. The models take text, image, video and audio as input, generate text (including code) and support a context window of up to 1 million tokens.

The series comprises three variants: the flagship MiMo-V2.6-Pro-RL (sparse Mixture-of-Experts, 1.02T total / 42B activated parameters, 384 routed experts with 8 active), the efficiency-balanced MiMo-V2.6-Flash-RL (309B / 15B activated) and the distilled MiMo-V2.6-Distill-Qwen-9B (9B, SFT on top of Qwen3.5-9B). The architecture combines a hybrid sliding-window-attention (SWA) backbone, a MiMo ViT vision encoder (681M), audio encoders and a 5-layer speculative decoder (Multi-Token Prediction).

The models are trained via scaled reinforcement learning (Group Relative Policy Optimization) in a single mixed RL run spanning coding, general agents, visual tasks and cybersecurity, complemented by on-policy distillation (MOPD2) and a self-improvement loop (Groupwise Reward Synthesis, Groupwise Advantage Redistribution). The models support tool calling and agentic operation.

Classification
MultimodalReasoning modelTool-using model
Family: MiMo
Access & deployment
APIDownloadHosted
CloudLocal
Weights: Open source
Key parameters
📏 Context: 1M tokenów
🧩 Parameters: 1,02 bln / 42 mld aktywnych (Pro-RL); 309 mld / 15 mld (Flash-RL); 9 mld (Distill-Qwen-9B)
Tools · ✓ Fine-tuning
📥 Input: text, image, video, audio

Technical specification

Context window
1M tokenów
tokens
Parameters
1,02 bln / 42 mld aktywnych (Pro-RL); 309 mld / 15 mld (Flash-RL); 9 mld (Distill-Qwen-9B)
parameters
License
MIT
Hardware requirements
Large variants (Pro-RL, Flash-RL) require multi-GPU deployment with tensor parallelism; SGLang or vLLM recommended. The 9B Distill-Qwen-9B variant can run on a single GPU.
Features:Tool useFine-tuning
Modalities
⬇ Input
textimagevideoaudio
⬆ Output
textcode

Capabilities and applications

Native model capabilities
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Agentic coding
Multi-hour, multi-step programming tasks performed autonomously by the model: cloning a repository, running tests, iterating on fixes, integrating with CLI tools. Characteristic of Codex variants (GPT-5.1-Codex-Mini, Codex-Max).
Category: coding
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Multimodal understanding
Category: multimodal
Image understanding
Analysing and interpreting the content of images.
Category: vision
Video understanding
The model's ability to analyse and interpret video content — recognising actions, motion, events and relationships between objects over time.
Category: video
Audio understanding
Category: audio
Cybersecurity
The model ability to perform computer-security tasks: vulnerability analysis, proof-of-concept exploit generation, patching, and cybersecurity question answering.
Category: other
Computer use
The model's ability to operate a computer interface by interpreting screenshots and generating actions such as clicks, typing, and navigating applications.
Category: planning

Benchmark results

7 benchmarks
DeepSWE v1.1
Code Agent (MiMo-V2.6 Pro)
71.9
📅 21 Sept 2026📄 Raport techniczny MiMo-V2.6
Terminal Bench 2.1
General Agent (MiMo-V2.6 Pro)
89.9
📅 21 Sept 2026📄 Raport techniczny MiMo-V2.6
Toolathlon-Verified
General Agent (MiMo-V2.6 Pro)
76.9
📅 21 Sept 2026📄 Raport techniczny MiMo-V2.6
OSWorld
General Agent (MiMo-V2.6 Pro)
82.0
📅 21 Sept 2026📄 Raport techniczny MiMo-V2.6
GDPval-AA
General Agent (MiMo-V2.6 Pro)
1673
📅 21 Sept 2026📄 Raport techniczny MiMo-V2.6
Agents' Last Exam
General Agent (MiMo-V2.6 Pro)
31.6
📅 21 Sept 2026📄 Raport techniczny MiMo-V2.6
CyberGym
Cybersecurity (MiMo-V2.6 Pro)
94.0
📅 21 Sept 2026📄 Raport techniczny MiMo-V2.6