Robots Atlas>ROBOTS ATLAS
Llama 3.2 3B

Llama 3.2 3B

3.2 3Bย ยทย Family: Llama
Meta lightweight text-only language model (3.21B parameters) with a 128K-token context window, optimized for edge and mobile devices.
โœ“ Activeโœ“ Public accessโš– Open weightsLLM๐Ÿ“ Llama
Context window
128K
tokens
Parameters
3.21B
parameters
Release date
25 September 2024
Access:DownloadAPIDeployment:๐Ÿ’ป Localโ˜ Cloud๐Ÿ“ฑ On-device

Overview

Llama 3.2 3B is a lightweight, text-only language model developed by Meta and released on September 25, 2024. It has about 3.21 billion parameters and a 128K-token context window. It was created through structured pruning of the Llama 3.1 8B model and knowledge distillation from the larger Llama 3.1 8B and 70B models.

It uses an auto-regressive transformer architecture with Grouped-Query Attention (GQA) and shared embeddings. Post-training included Supervised Fine-Tuning (SFT), Rejection Sampling and Direct Preference Optimization (DPO).

It officially supports eight languages (English, German, French, Italian, Portuguese, Hindi, Spanish and Thai) and tool calling. It is designed for edge and mobile devices - Meta released it in cooperation with Qualcomm, MediaTek and Arm. The model knowledge cutoff is December 2023. It is distributed under the Llama 3.2 Community License.

Classification
LLM
Family: Llama
Access & deployment
DownloadAPI
LocalCloudOn-device
Weights: Open weights
Key parameters
๐Ÿ“ Context: 128K
๐Ÿงฉ Parameters: 3.21B
โœ“ Toolsย ยทย โœ“ Fine-tuning
๐Ÿ“ฅ Input: text

Technical specification

Context window
128K
tokens
Parameters
3.21B
parameters
Knowledge cutoff
1 Dec 2023
Knowledge boundary
License
Llama 3.2 Community License
Hardware requirements
Optimized for edge and mobile devices (Qualcomm, MediaTek, Arm processors). BF16 weights; quantized variants also available.
Features:โœ“ Tool useโœ“ Fine-tuning
Modalities
โฌ‡ Input
text
โฌ† Output
textcode

Capabilities and applications

Native model capabilities
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Instruction following
Precisely following instructions contained in the prompt: response format, length, style, constraints (e.g. 'reply in six words'). GPT-5.1 significantly improved this capability compared to GPT-5.
Category: language
Long context
Support for large context windows โ€” tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Language modeling
Ability to predict subsequent tokens and generate coherent natural-language text based on the preceding context.
Category: language
Retrieval-Augmented Generation (RAG)
The model's ability to generate answers grounded in retrieved documents/data, with source attribution (citations).
Category: language

Benchmark results

3 benchmarks
MMLU
macro_avg/acc ยท 5-shot
63.4%
๐Ÿ“„ Hugging Face model card (Meta)
GSM8K
em_maj1@1 ยท 8-shot, CoT
77.7%
๐Ÿ“„ Hugging Face model card (Meta)
IFEval
average ยท instruction-tuned
77.4%
๐Ÿ“„ Hugging Face model card (Meta)

Technical architecture

Core Architecture
Model Form