Robots Atlas>ROBOTS ATLAS
Mistral Small 4

Mistral Small 4

119B-2603 (26.03)ย ยทย Family: Mistral
Open Mixture-of-Experts model from Mistral AI (119B params, ~6B active) unifying instruct and reasoning, with text+image input and a 256k context window.
โœ“ Activeโœ“ Public accessโš– Open sourceLLMMultimodalReasoning model๐Ÿ“ Mistral
Context window
256K
tokens
Parameters
119B (MoE, ~6B aktywnych)
parameters
Release date
16 March 2026
Access:APIDownloadHostedDeployment:โ˜ Cloud๐Ÿ’ป Local

Overview

Mistral Small 4 (codename 2603, repository Mistral-Small-4-119B-2603) is an open language model from Mistral AI released on 16 March 2026 under the Apache 2.0 license. It uses a Mixture-of-Experts architecture (128 experts, 4 active per token): with 119B total parameters it activates about 6B parameters per token.

The model supports text and image input, produces text, offers a 256k token context window, and provides native function calling and JSON output. It is a hybrid model unifying instruct behaviour with reasoning controlled by the reasoning_effort parameter. It is available via the Mistral API, in Mistral AI Studio, on Hugging Face, and across the NVIDIA ecosystem (build.nvidia.com, NIM) and frameworks including vLLM, llama.cpp, SGLang and Transformers.

Classification
LLMMultimodalReasoning model
Family: Mistral
Access & deployment
APIDownloadHosted
CloudLocal
Weights: Open source
Key parameters
๐Ÿ“ Context: 256K
๐Ÿงฉ Parameters: 119B (MoE, ~6B aktywnych)
โœ“ Toolsย ยทย โœ“ Fine-tuning
๐Ÿ“ฅ Input: text, image

Technical specification

Context window
256K
tokens
Parameters
119B (MoE, ~6B aktywnych)
parameters
License
Apache 2.0
Features:โœ“ Tool useโœ“ Fine-tuning
Modalities
โฌ‡ Input
textimage
โฌ† Output
text

Capabilities and applications

Native model capabilities
Function Calling
Category: planning
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Adaptive reasoning effort
The model decides how much 'thinking' to allocate to a given query: simple questions are answered quickly, complex problems receive more inference cycles. A GPT-5.1 feature (both Instant and Thinking) that shortens time on easy tasks and extends it for hard ones.
Category: reasoning
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Image understanding
Analysing and interpreting the content of images.
Category: vision
Multimodal understanding
Category: multimodal
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Long context
Support for large context windows โ€” tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Structured output
Producing data in structured formats such as JSON.
Category: structured_generation
Instruction following
Precisely following instructions contained in the prompt: response format, length, style, constraints (e.g. 'reply in six words'). GPT-5.1 significantly improved this capability compared to GPT-5.
Category: language

Benchmark results

1 benchmark
GPQA
accuracy
71.2%
๐Ÿ“… 16 Mar 2026๐Ÿ“„ Oficjalna karta modelu (Hugging Face)

Pricing

Technical architecture