
Perceptron AI's open foundation model for robot learning, uniting multimodal video understanding, embodied reasoning and robot control in one 36B sparse (MoE) model.
✓ Active✓ Public access⚖ Open sourceRobotics foundation modelVision-Language-Action modelMultimodal
Parameters
36B (sparse MoE)
parameters
Release date
26 August 2026
Access:DownloadDeployment:💻 Local📱 On-device
Overview
Classification
Robotics foundation modelVision-Language-Action modelMultimodal
Applications
Access & deployment
Download
LocalOn-device
Weights: Open source
Key parameters
🧩 Parameters: 36B (sparse MoE)
✓ Fine-tuning
📥 Input: text, image, video, robot state data
Robotics
Robot controlRobot manipulationVisual groundingSpatial reasoningSpatial predictionEmbodied task planningScene understandingEnvironment modeling
Technical specification
Parameters
36B (sparse MoE)
parameters
License
Apache 2.0
Features:✓ Fine-tuning
Modalities
⬇ Input
textimagevideorobot_state_data
⬆ Output
textstructured_datarobot_actionsmotion_trajectories
Capabilities and applications
Native model capabilities
Video understanding
The model's ability to analyse and interpret video content — recognising actions, motion, events and relationships between objects over time.
Category: video
Multimodal understanding
Category: multimodal
Image understanding
Analysing and interpreting the content of images.
Category: vision
Object tracking (video)
The ability to track selected objects across consecutive video frames, maintaining their masks/identity despite motion, occlusion and appearance changes.
Category: vision
Vision-language-action grounding
The ability of a VLA model to ground visual perception and a language instruction into a concrete physical robot action. The model understands the scene and intent, then generates an executable action sequence, closing the loop from observation to motion.
Category: robotics
Action conditioning
Controlling model generation via action signals (camera, robot pose, commands, speech) rather than text prompts alone.
Category: multimodal
Cross-embodiment transfer
The ability of a single model to control robots with different morphologies (humanoids, dual-arm rigs, mobile platforms) without training a separate model per platform. Intelligence is decoupled from embodiment, so the same policy runs on hardware with different kinematics and dynamics.
Category: robotics
Real-time inference
The model's ability to generate responses with very low latency (>1000 tokens/sec) on specialized inference hardware (e.g. Cerebras WSE), enabling interactive, turn-by-turn collaboration with a human.
Category: coding
Robotics
Robot controlRobot manipulationVisual groundingSpatial reasoningSpatial predictionEmbodied task planningScene understandingEnvironment modeling
Application domains