Robots Atlas>ROBOTS ATLAS
WALL-A

WALL-A

WALL-A · Family: WALL
X Square Robot’s embodied operation model (WALL series): end-to-end, tens-of-billions parameters, World Unified Model architecture and X-Tokenizer. Robotic manipulation.
✓ Active🏢 EnterpriseRobotics foundation modelVision-Language-Action model📁 WALL
Parameters
百亿级 (skala dziesiątek miliardów, wg producenta)
parameters
Deployment:📱 On-device💻 Local

Overview

WALL-A is an embodied “operation model” (操作大模型) developed by X Square Robot (自变量机器人, Shenzhen), serving as the core of the WALL family of foundation models. The company describes it as one of China’s earliest fully end-to-end implementations of embodied AI, integrating perception, understanding and action control in a single model.

The model uses an end-to-end architecture at a tens-of-billions parameter scale (百亿级, per the vendor’s materials) and leverages the World Unified Model architecture together with the X-Tokenizer action tokenizer. It takes vision (RGB-D, 3D-ToF sensors), the robot’s proprioceptive data and natural-language instructions as input, and produces action-control sequences as output.

WALL-A supports few-shot learning across tasks, manipulation of difficult objects (friction, fluids, deformable objects), force control and precise positioning, embodied chain-of-thought reasoning, and long operation sequences with sub-millimeter precision on high-degree-of-freedom systems.

The model is deployed in X Square Robot’s robots, including the Quantum X1 Pro (wheeled dual-arm), the Quantum X2 (wheeled humanoid) and the five-finger ArtiXon Hand. It is a proprietary (closed) solution available through the company’s products rather than as a separately downloadable model.

Classification
Robotics foundation modelVision-Language-Action model
Family: WALL
Access & deployment
On-deviceLocal
Weights: Closed
Key parameters
🧩 Parameters: 百亿级 (skala dziesiątek miliardów, wg producenta)
📥 Input: text, image, depth, robot sensors
Robotics
Robot manipulationDexterous manipulationBimanual manipulationEmbodied task planningRobot controlSpatial reasoningObject affordance understanding

Technical specification

Parameters
百亿级 (skala dziesiątek miliardów, wg producenta)
parameters
License
Proprietary
Hardware requirements
Deployed on X Square Robot’s robots (Quantum X1 Pro, Quantum X2, ArtiXon Hand).
Modalities
⬇ Input
textimagedepthrobot_sensorsrobot_state_data
⬆ Output
robot_actionsmanipulator_controlrobot_commands

Capabilities and applications

Native model capabilities
Advanced reasoning
The ability to perform multi-step, structured reasoning: analysing problems, planning steps, and drawing conclusions from hypotheses. Reasoning-first models (e.g. GPT-5.1 Thinking) dedicate a portion of inference to chains of thought before responding.
Category: reasoning
Multi-step reasoning
Carrying out multi-step chains of reasoning across long, complex tasks.
Category: reasoning
Multimodal understanding
Category: multimodal
Planning
Forming and executing action plans for complex tasks.
Category: planning
Instruction following
Precisely following instructions contained in the prompt: response format, length, style, constraints (e.g. 'reply in six words'). GPT-5.1 significantly improved this capability compared to GPT-5.
Category: language
Robotics
Robot manipulationDexterous manipulationBimanual manipulationEmbodied task planningRobot controlSpatial reasoningObject affordance understanding
Application domains

Technical architecture