Robots Atlas>ROBOTS ATLAS
Gemma-2-9B-it

Gemma-2-9B-it

2 · 9B · Instruct · Family: Gemma
An open, instruction-tuned 9B language model from Google DeepMind. Text-to-text, 8192 context, alternating local/global attention, GQA, knowledge distillation. Gemma license.
✓ Active✓ Public access⚖ Open weightsLLM📁 Gemma
Context window
8K
tokens
Parameters
9B
parameters
Release date
27 June 2024
Access:DownloadHostedDeployment:💻 Local☁ Cloud📱 On-device

Overview

Gemma-2-9B-it is an open, instruction-tuned large language model (LLM) with 9 billion parameters, developed by Google DeepMind and released in June 2024. It is a decoder-only, text-to-text model designed for tasks such as question answering, summarization, reasoning, and code generation.

The Gemma 2 architecture combines alternating local (sliding-window) and global attention layers, Grouped-Query Attention (GQA), and logit soft-capping in the attention and final layers. The 9B variant was trained using knowledge distillation from a larger teacher model. The model uses an 8192-token context window.

The model was trained on 8 trillion tokens using JAX and ML Pathways on TPU hardware (TPUv5p). The -it variant underwent supervised instruction tuning (SFT) and reinforcement learning from human feedback (RLHF). Weights are released under the Gemma license (requires acceptance of Google's terms); the model can be downloaded and run locally or in the cloud.

Classification
LLM
Family: Gemma
Access & deployment
DownloadHosted
LocalCloudOn-device
Weights: Open weights
Key parameters
📏 Context: 8K
🧩 Parameters: 9B
✓ Fine-tuning
📥 Input: text

Technical specification

Context window
8K
tokens
Parameters
9B
parameters
License
Gemma
Hardware requirements
Native weights in bfloat16. Running the 9B variant is recommended on a GPU/TPU with sufficient memory; quantized variants are available for local use.
Features:✓ Fine-tuning
Modalities
⬇ Input
text
⬆ Output
textcode

Capabilities and applications

Native model capabilities
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Instruction following
Precisely following instructions contained in the prompt: response format, length, style, constraints (e.g. 'reply in six words'). GPT-5.1 significantly improved this capability compared to GPT-5.
Category: language
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding

Benchmark results

9 benchmarks
MMLU
accuracy · 5-shot
71.3%
📄 technical_report
Score for the Gemma 2 9B pretrained (base) variant.
HumanEval
pass@1 · 0-shot
40.2%
📄 technical_report
GSM8K
accuracy · 5-shot, maj@1
68.6%
📄 technical_report
MATH
accuracy · 4-shot
36.6%
📄 technical_report
ARC-c
accuracy · 25-shot
68.4%
📄 technical_report
HellaSwag
accuracy · 10-shot
81.9%
📄 technical_report
Winogrande
accuracy · 5-shot
80.6%
📄 technical_report
BoolQ
accuracy · 0-shot
84.2%
📄 technical_report
PIQA
accuracy · 0-shot
81.7%
📄 technical_report

Technical architecture

Core Architecture
Model Form