Robots Atlas>ROBOTS ATLAS
ERNIE X1

ERNIE X1

ERNIE X1 · Family: ERNIE
Baidu’s reasoning (“deep thinking”) model, announced 16 March 2025: FlashMask + heterogeneous multimodal MoE architecture; powers the Ernie Bot assistant.
✓ Active✓ Public accessLLMReasoning modelMultimodal📁 ERNIE
Release date
16 March 2025
Access:APIHostedDeployment:☁ Cloud

Overview

ERNIE X1 is a reasoning (“deep thinking”) model developed by Baidu, announced on 16 March 2025 alongside the general-purpose ERNIE 4.5 model. It is a specialized model focused on complex reasoning and is part of the ERNIE family (Chinese 文心, Wenxin) that powers the Ernie Bot assistant.

Its performance improvements were achieved through new technologies, including “FlashMask” dynamic attention masking and a heterogeneous multimodal mixture-of-experts architecture. Thanks to the multimodal architecture, the model can process different data types as part of its reasoning process.

In April 2025, Baidu released Turbo variants — ERNIE 4.5 Turbo and ERNIE X1 Turbo — optimized for faster responses and lower operational costs.

The model is available through the Ernie Bot assistant (yiyan.baidu.com), which Baidu made free of charge, and through Baidu’s cloud platform (Qianfan). ERNIE X1 is a proprietary model — its weights have not been released (unlike some later ERNIE models).

Classification
LLMReasoning modelMultimodal
Family: ERNIE
Access & deployment
APIHosted
Cloud
Weights: Closed
Key parameters
📥 Input: text, image

Technical specification

License
Proprietary
Hardware requirements
A proprietary, hosted (cloud) model — delivered as a service via the Ernie Bot assistant and Baidu AI Cloud (Qianfan) platform. No local deployment (closed weights).
Modalities
⬇ Input
textimage
⬆ Output
text

Capabilities and applications

Native model capabilities
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Advanced reasoning
The ability to perform multi-step, structured reasoning: analysing problems, planning steps, and drawing conclusions from hypotheses. Reasoning-first models (e.g. GPT-5.1 Thinking) dedicate a portion of inference to chains of thought before responding.
Category: reasoning
Multi-step reasoning
Carrying out multi-step chains of reasoning across long, complex tasks.
Category: reasoning
Extended thinking mode
A reasoning-model variant with a larger inference budget: more thinking cycles, higher answer precision at the cost of response time. Choice between 'standard' and 'extended' thinking is left to the user (e.g. the selector in GPT-5.2 Pro).
Category: reasoning
Mathematical reasoning
The model's ability to solve mathematical tasks requiring multi-step reasoning — equations, proofs, combinatorics, geometry, calculus and competition-level problems.
Category: reasoning
Image understanding
Analysing and interpreting the content of images.
Category: vision
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Instruction following
Precisely following instructions contained in the prompt: response format, length, style, constraints (e.g. 'reply in six words'). GPT-5.1 significantly improved this capability compared to GPT-5.
Category: language
Natural conversation
Conducting a conversation with a tone close to human: a warmer voice, empathy in emotional responses, humour, and avoiding the stiff 'AI assistant jargon'. Introduced as a deliberate improvement in GPT-5.1 Instant.
Category: language

Technical architecture

Core Architecture
Training Techniques