Robots Atlas>ROBOTS ATLAS
Mistral 7B Instruct

Mistral 7B Instruct

Mistral-7B-Instruct-v0.3 (7.3B) · Family: Mistral
Open (Apache 2.0) 7.3B-parameter instruction-tuned language model from Mistral AI; 32k context window, v3 tokenizer and function calling (v0.3).
✓ Active✓ Public access⚖ Open sourceLLMTool-using model📁 Mistral
Context window
32K
tokens
Parameters
7.3B
parameters
Release date
27 September 2023
Access:DownloadAPIDeployment:💻 Local☁ Cloud

Overview

Mistral 7B Instruct is the instruction-tuned variant of the open Mistral 7B language model developed by the French company Mistral AI. The base model has 7.3B parameters and uses a Transformer architecture with Grouped-Query Attention (GQA) for faster inference and — in the original v0.1 — Sliding Window Attention (SWA) with a 4096 window.

In versions 0.2 and 0.3 Sliding Window Attention was disabled and the context window was extended to 32,768 tokens (rope-theta = 1e6). Version 0.3 additionally extends the vocabulary to 32,768 tokens, introduces the v3 tokenizer and adds function calling support.

The model is released under the Apache 2.0 license — with open weights available for download (Hugging Face) and through the Mistral API (la Plateforme).

Classification
LLMTool-using model
Family: Mistral
Access & deployment
DownloadAPI
LocalCloud
Weights: Open source
Key parameters
📏 Context: 32K
🧩 Parameters: 7.3B
✓ Tools · ✓ Fine-tuning
📥 Input: text

Technical specification

Context window
32K
tokens
Parameters
7.3B
parameters
License
Apache 2.0
Hardware requirements
Open weights (Apache 2.0) on Hugging Face. At 7.3B parameters the model runs locally on a single consumer-grade GPU (as little as ~8 GB VRAM when quantized). Also available via the Mistral API (la Plateforme).
Features:✓ Tool use✓ Fine-tuning
Modalities
⬇ Input
text
⬆ Output
textcode

Benchmark results

8 benchmarks
MMLU
accuracy · Mistral 7B base model, 5-shot
60.1%
📄 Jiang et al., Mistral 7B (arXiv:2310.06825)
Score for the base Mistral 7B model from the official paper; the Instruct variant is built on this base.
HellaSwag
accuracy · Mistral 7B base model
81.3%
📄 Jiang et al., Mistral 7B (arXiv:2310.06825)
Score for the base Mistral 7B model from the official paper.
WinoGrande
accuracy · Mistral 7B base model
75.3%
📄 Jiang et al., Mistral 7B (arXiv:2310.06825)
Score for the base Mistral 7B model from the official paper.
ARC-Challenge
accuracy · Mistral 7B base model
55.5%
📄 Jiang et al., Mistral 7B (arXiv:2310.06825)
Score for the base Mistral 7B model from the official paper.
HumanEval
pass@1 · Mistral 7B base model
30.5%
📄 Jiang et al., Mistral 7B (arXiv:2310.06825)
Score for the base Mistral 7B model from the official paper.
MBPP
pass@1 · Mistral 7B base model
47.5%
📄 Jiang et al., Mistral 7B (arXiv:2310.06825)
Score for the base Mistral 7B model from the official paper.
MATH
accuracy · Mistral 7B base model, 4-shot
13.1%
📄 Jiang et al., Mistral 7B (arXiv:2310.06825)
Score for the base Mistral 7B model from the official paper.
GSM8K
accuracy · Mistral 7B base model, 8-shot maj@8
52.2%
📄 Jiang et al., Mistral 7B (arXiv:2310.06825)
Score for the base Mistral 7B model from the official paper.

Technical architecture

Core Architecture
Model Form
Training Techniques