Robots Atlas>ROBOTS ATLAS
MiniMax-M3-MXFP8

MiniMax-M3-MXFP8

M3 (MXFP8)
MXFP8 (microscaling FP8) quantized variant of MiniMax-M3. Natively multimodal MoE model (~428B total, ~23B active) with a 1M-token context.
✓ Active✓ Public access⚖ Open weightsLLMMultimodalReasoning modelTool-using model
Context window
1M
tokens
Parameters
428B (23B active)
parameters
Access:APIDownloadHostedDeployment:☁ Cloud💻 Local

Overview

MiniMax-M3-MXFP8 is a variant of MiniMax-M3, developed by MiniMax, quantized to the MXFP8 (microscaling FP8) format. MXFP8 quantization reduces weight size and memory footprint while preserving model quality, easing deployment on FP8-capable accelerators.

The base MiniMax-M3 is a natively multimodal Mixture-of-Experts model with about 428B parameters, of which about 23B are activated per token. It supports up to a 1M-token context window via MiniMax Sparse Attention (MSA). It accepts text, image and video as input and produces text.

The model supports agentic use, coding and reasoning modes controlled by the thinking parameter (enabled, adaptive, disabled). Weights are released on Hugging Face under the MiniMax Community license; recommended inference frameworks are SGLang, vLLM and Transformers. The model is also available via the MiniMax API and Agent.

Classification
LLMMultimodalReasoning modelTool-using model
Access & deployment
APIDownloadHosted
CloudLocal
Weights: Open weights
Key parameters
📏 Context: 1M
🧩 Parameters: 428B (23B active)
Tools
📥 Input: text, image, video

Technical specification

Context window
1M
tokens
Parameters
428B (23B active)
parameters
License
MiniMax Community License
Hardware requirements
The MXFP8 (microscaling FP8) variant reduces the memory footprint versus the BF16 original. Recommended inference frameworks: SGLang, vLLM, Transformers.
Features:Tool use
Modalities
⬇ Input
textimagevideo
⬆ Output
text

Capabilities and applications

Native model capabilities
Long context
Processing very long inputs (tens to hundreds of thousands of tokens) while maintaining coherence.
Category: language
Coding
Generating, completing, explaining and debugging code across multiple programming languages.
Category: coding
Reasoning
The model's ability to perform multi-step logical inference, solve complex problems and decompose tasks into steps.
Category: reasoning
Interleaved Multimodal Input
Category: reasoning
Function Calling
Category: planning

Pricing

Technical architecture