Robots Atlas>ROBOTS ATLAS
Mercury 2

Mercury 2

2
Diffusion-based reasoning LLM (dLLM) from Inception Labs that generates tokens in parallel. 128K context window, native tool use, public pricing.
✓ Active✓ Public accessReasoning modelLLM
Context window
128K
tokens
Release date
24 February 2026
Access:APIHostedDeployment:☁ Cloud

Overview

Mercury 2 is a language model based on a diffusion architecture (diffusion LLM, dLLM) developed by Inception Labs. Unlike classic autoregressive models that generate text token by token from left to right, Mercury 2 produces its answer in parallel, progressively refining whole blocks of text ("coarse-to-fine"). This allows for much faster generation while maintaining high quality.

The model combines tunable reasoning with a 128K context window, native tool use and schema-aligned JSON output; its API is OpenAI-compatible. According to Inception Labs, Mercury 2 reaches around 1,009 tokens per second on NVIDIA Blackwell GPUs and a time-to-first-token (TTFT) below 300 ms on standard NVIDIA GPUs.

Mercury 2 belongs to the Mercury family, described by the company as the first commercially available family of diffusion large language models. It was released on 24 February 2026 and is available via API (platform.inceptionlabs.ai) and a chat interface (chat.inceptionlabs.ai).

Classification
Reasoning modelLLM
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📏 Context: 128K
✓ Tools
📥 Input: text

Technical specification

Context window
128K
tokens
Features:✓ Tool use
Modalities
⬇ Input
text
⬆ Output
textcode

Capabilities and applications

Native model capabilities
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Structured output
Producing data in structured formats such as JSON.
Category: structured_generation
Real-time inference
The model's ability to generate responses with very low latency (>1000 tokens/sec) on specialized inference hardware (e.g. Cerebras WSE), enabling interactive, turn-by-turn collaboration with a human.
Category: coding
Adaptive reasoning effort
The model decides how much 'thinking' to allocate to a given query: simple questions are answered quickly, complex problems receive more inference cycles. A GPT-5.1 feature (both Instant and Thinking) that shortens time on easy tasks and extends it for hard ones.
Category: reasoning

Pricing

Technical architecture

Core Architecture