Vision-language-action model from the UBTECH Thinker family: combines visual perception, language instructions and action generation. Applied to industrial manipulation tasks.
Deployment:📱 On-device☁ Cloud
Overview
Applications
Access & deployment
On-deviceCloud
Weights: Open source
Key parameters
📥 Input: image, text, robot sensors
Robotics
Robot manipulationRobot controlVisual groundingObject affordance understandingMotion planning
Technical specification
License
Open source (deklaracja UBTECH; brak potwierdzonego publicznego repozytorium)
Hardware requirements
Not publicly disclosed. A VLA execution model in the action layer of UBTECH humanoids (on-device and cloud inference).
Modalities
⬇ Input
imagetextrobot_sensors
⬆ Output
robot_actionsmanipulator_controlmotion_trajectories
Capabilities and applications
Native model capabilities
Vision-language-action grounding
The ability of a VLA model to ground visual perception and a language instruction into a concrete physical robot action. The model understands the scene and intent, then generates an executable action sequence, closing the loop from observation to motion.
Category: robotics
Action conditioning
Controlling model generation via action signals (camera, robot pose, commands, speech) rather than text prompts alone.
Category: multimodal
Planning
Forming and executing action plans for complex tasks.
Category: planning
Multimodal understanding
Category: multimodal
Robotics
Robot manipulationRobot controlVisual groundingObject affordance understandingMotion planning
Application domains
Technical architecture
Core Architecture
Model Form
