Robots Atlas>ROBOTS ATLAS
Inference

Top-p / Nucleus Sampling

2019ActivePublished
Key innovation
Instead of truncating to a fixed number of tokens (top-k), nucleus sampling dynamically selects the minimal set of tokens whose cumulative probability ≥ p, providing better diversity control across variable distributions.
Category
Inference
Abstraction level
Building block
Operation level
Inference
Use cases
Text generation in ChatGPT, Claude, GeminiCreative writing with controlled diversityCode generation balancing quality and variationDefault sampling method in most LLM APIsConversational chatbots — avoiding repetition

How it works

After computing softmax, the model sorts tokens in descending probability order. It then selects the minimal token prefix whose cumulative probability ≥ p (e.g., p=0.9). Only tokens within this "nucleus" are considered for sampling. Tokens outside the nucleus receive probability 0.

Problem solved

Greedy decoding yields monotonous text, while temperature-only sampling can generate incoherent tokens. Top-p balances creativity and quality by dynamically constraining the sampling space to the "nucleus" of the distribution.

Implementation

Implementation pitfalls
Top-p and top-k together can over-constrain the spaceMedium

Using top-p=0.9 and top-k=50 simultaneously: first top-k reduces to 50 tokens, then top-p to ~90% mass — resulting in doubly truncated space that may eliminate valid tokens.

High p for flat distributions = near-random samplingMedium

With p=0.99 and a flat distribution (e.g. 1000 tokens with similar probability) the nucleus contains ~990 tokens — practically random sampling without filtering.

Hyperparameters (configurable axes)

p (nucleus threshold)Critical

Cumulative probability threshold. Higher p → more tokens in nucleus → more diversity.

0.9
0.95
1.0 (wyłącza filtrowanie)

Execution paradigm

Primary mode
Conditional
Activation pattern
Input dependent

Parallelism

Parallelism level
Sequential
Scope
Inference