Robots Atlas>ROBOTS ATLAS
Llama 2 7B Chat

Llama 2 7B Chat

7B Chat · Family: Llama
Meta 7-billion-parameter language model, Chat variant fine-tuned with RLHF for dialogue. 4,096-token context window, open weights under the Llama 2 Community License.
✓ Active✓ Public access⚖ Open weightsLLM📁 Llama
Context window
4K
tokens
Parameters
7B
parameters
Release date
18 July 2023
Access:DownloadHostedAPIDeployment:💻 Local☁ Cloud

Overview

Llama 2 7B Chat is the smallest conversational model in the Llama 2 family developed by Meta and released in July 2023. It is built on an autoregressive Transformer architecture with 7 billion parameters, trained on roughly 2 trillion tokens of text data, with a knowledge cutoff of September 2022.

The Chat variant was produced by fine-tuning the base model: supervised fine-tuning on instructions (SFT) and reinforcement learning from human feedback (RLHF), optimizing the model for dialogue, helpfulness and safety. The model takes and generates text only, with a 4,096-token context window.

The model weights are open (open weights) and distributed under the Llama 2 Community License, which permits commercial use with restrictions (including a 700M monthly active users threshold) and is not an OSI-approved open-source license. Unlike the 34B and 70B variants, the 7B variant does not use Grouped-Query Attention.

Classification
LLM
Family: Llama
Access & deployment
DownloadHostedAPI
LocalCloud
Weights: Open weights
Key parameters
📏 Context: 4K
🧩 Parameters: 7B
✓ Fine-tuning
📥 Input: text

Technical specification

Context window
4K
tokens
Parameters
7B
parameters
Knowledge cutoff
1 Sept 2022
Knowledge boundary
License
Llama 2 Community License
Hardware requirements
The FP16 variant requires roughly 14 GB of memory, allowing it to run on a single GPU (e.g. 16 GB). After 4-bit quantization the requirement drops to about 4-5 GB.
Features:✓ Fine-tuning
Modalities
⬇ Input
text
⬆ Output
textcode

Capabilities and applications

Native model capabilities
Natural conversation
Conducting a conversation with a tone close to human: a warmer voice, empathy in emotional responses, humour, and avoiding the stiff 'AI assistant jargon'. Introduced as a deliberate improvement in GPT-5.1 Instant.
Category: language
Language modeling
Ability to predict subsequent tokens and generate coherent natural-language text based on the preceding context.
Category: language
Instruction following
Precisely following instructions contained in the prompt: response format, length, style, constraints (e.g. 'reply in six words'). GPT-5.1 significantly improved this capability compared to GPT-5.
Category: language
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding

Benchmark results

4 benchmarks
MMLU
accuracy · 5-shot
45.3%
📄 paper
Llama 2 7B base model (pretrained).
GSM8K
accuracy · 8-shot
14.6%
📄 paper
Llama 2 7B base model (pretrained).
TruthfulQA
truthful & informative · Llama-2-Chat 7B variant (safety fine-tuned).
57.04%
📄 Karta modelu Meta (Hugging Face)
Higher is better (more truthful and informative answers).
ToxiGen
toxic generations · Llama-2-Chat 7B variant; lower is better.
0.00%
📄 Karta modelu Meta (Hugging Face)
Percentage of toxic generations.

Technical architecture

Model Form
Training Techniques