Robots Atlas>ROBOTS ATLAS
UBTECH Thinker-VLA

UBTECH Thinker-VLA

Thinker-VLA 1.0 · Family: Thinker
Vision-language-action model from the UBTECH Thinker family: combines visual perception, language instructions and action generation. Applied to industrial manipulation tasks.
✓ Active✓ Public access⚖ Open sourceVision-Language-Action model📁 Thinker
Deployment:📱 On-device☁ Cloud

Overview

UBTECH Thinker-VLA is a vision-language-action (VLA) model from the Thinker family by UBTECH Robotics. It is the third, execution tier of the Thinker stack: it combines visual perception, natural-language instructions and action generation, adapting the robot movements as objects and surroundings change.

Applications and performance

The model is used in industrial handling and loading tasks. According to UBTECH it improves inference efficiency in industrial scenarios by 176%. Architecture details, parameter counts and license have not been publicly disclosed.

Classification
Vision-Language-Action model
Family: Thinker
Access & deployment
On-deviceCloud
Weights: Open source
Key parameters
📥 Input: image, text, robot sensors
Robotics
Robot manipulationRobot controlVisual groundingObject affordance understandingMotion planning

Technical specification

License
Open source (deklaracja UBTECH; brak potwierdzonego publicznego repozytorium)
Hardware requirements
Not publicly disclosed. A VLA execution model in the action layer of UBTECH humanoids (on-device and cloud inference).
Modalities
⬇ Input
imagetextrobot_sensors
⬆ Output
robot_actionsmanipulator_controlmotion_trajectories

Capabilities and applications

Native model capabilities
Vision-language-action grounding
The ability of a VLA model to ground visual perception and a language instruction into a concrete physical robot action. The model understands the scene and intent, then generates an executable action sequence, closing the loop from observation to motion.
Category: robotics
Action conditioning
Controlling model generation via action signals (camera, robot pose, commands, speech) rather than text prompts alone.
Category: multimodal
Planning
Forming and executing action plans for complex tasks.
Category: planning
Multimodal understanding
Category: multimodal
Robotics
Robot manipulationRobot controlVisual groundingObject affordance understandingMotion planning

Technical architecture