Robots Atlas>ROBOTS ATLAS
Reasoning

Reasoning Effort / Reasoning Levels

2024ActiveUpdated: 14 August 2026Published
Key innovation
Exposes test-time compute as a tunable per-request parameter, letting the same model deliberately trade latency and cost for accuracy.
Category
Reasoning
Abstraction level
Primitive
Operation level
InferenceServing
Use cases
Hard math and logic tasks requiring multi-step reasoningAgentic coding and debuggingReducing cost and latency on simple queries (low effort)Tuning the quality/cost trade-off in production APIsTasks requiring self-verification and planning

How it works

The developer passes an effort level (e.g. low/medium/high) or a thinking-token budget in the API request. The model controls the length of its internal reasoning chain: at low effort it emits a short thinking trace (or skips it), at high effort a longer one, allowing more inference steps, self-verification and solution exploration before the final output. Because thinking tokens are generated autoregressively and are usually billed, higher effort directly increases latency and cost. The exact semantics (discrete levels vs. token budget, whether the thinking is visible) differ across providers.

Problem solved

Reasoning models perform better when they 'think longer', but a fixed, maximal amount of reasoning is slow and costly for simple queries. Reasoning Effort addresses the lack of control over the cost/latency/quality trade-off by putting it in the developer's hands per request.

Key mechanisms

A request parameter (e.g. reasoning_effort) or a thinking-token budget
Controlling the length of the reasoning chain (chain-of-thought)
Scaling inference-time compute (test-time compute)
Billing thinking tokens as cost

Strengths & limitations

Strengths
โœ“Flexible cost/latency/quality trade-off without switching models
โœ“Better results on hard tasks at high effort
โœ“Savings at low effort for simple queries
โœ“Easy to use (a single API parameter)
Limitations
โœ—Diminishing returns above a certain effort level
โœ—Risk of 'over-thinking' and wasting cost on easy tasks
โœ—Inconsistent semantics across providers (levels vs. token budget)
โœ—Hard to pick the right level without measurement
โœ—Higher latency hampers interactive use

Components

Effort parameterInput control signal

An API request field (e.g. reasoning_effort = low/medium/high) or a thinking-token budget controlling the amount of reasoning.

Reasoning-chain generatorExecutes test-time compute

The model mechanism that generates internal thinking tokens, whose length scales with the requested effort.

Implementation

Implementation pitfalls
Over-thinking on easy tasksMedium

High effort on simple queries wastes time and tokens with no quality gain.

Fix:Pick the lowest effort that meets the required quality; consider difficulty-based routing.
Uncontrolled cost/latencyHigh

Thinking tokens are usually billed and lengthen responses; high effort can multiply cost.

Fix:Set budget caps, monitor thinking tokens, and test levels on representative traffic.
Inconsistent semantics across providersMedium

Discrete levels vs. token budgets, differing defaults and thinking visibility hurt portability.

Fix:Abstract the effort configuration behind a common interface and map it per provider.

Evolution

Original paper ยท 2024 ยท OpenAI (blog) ยท OpenAI
Learning to Reason with LLMs (OpenAI o1)
OpenAI
2024
OpenAI o1 and the reasoning_effort parameter (low/medium/high)
Inflection point

o1 models introduce test-time compute as a controllable reasoning effort in the API.

2025
Anthropic 'extended thinking' with a token budget

Claude exposes a thinking-token budget as a continuous effort control.

2025
GPT-5 adds a 'minimal' level

Extends the effort scale with a cheapest minimal-reasoning mode.

2025
Google Gemini 3 โ€” thinking levels

Gemini exposes discrete thinking levels (low/high) alongside the earlier budget.

Hyperparameters (configurable axes)

Effort levelCritical

Discrete reasoning-effort level.

minimalGPT-5: minimal reasoning, fastest/cheapest.
lowShort reasoning.
mediumDefault trade-off.
highDeepest reasoning, slowest/most expensive.
Thinking token budgetHigh

An alternative, continuous control: the maximum number of tokens allotted to internal thinking (e.g. Anthropic budget_tokens, Gemini thinking budget).

1024โ€“32768+Typical thinking-token budget ranges.

Computational complexity

Time complexity: O(b).

Compute bottleneck

Sequential decoding of thinking tokens

The internal reasoning chain is generated autoregressively token by token, making latency the main constraint at high effort.

Execution paradigm

Primary mode
Conditional

The amount of compute depends on the requested effort (and, with adaptive thinking, on query difficulty).

Activation pattern
Input dependent

Parallelism

Parallelism level
Sequential

Longer thinking means a longer token sequence generated serially; parallelism applies more to 'parallel test-time compute' (many runs at once) than to a single chain.

Scope
InferenceAcross tokens

Hardware requirements

Primary

Generating long thinking traces is intensive LLM decoding on GPUs.

Good fit

Reasoning models are also served on TPUs (e.g. Gemini).

Possible

It is a control parameter, independent of specific hardware.