About
GEN is a family of embodied foundation models (Vision-Language-Action) developed by Generalist AI, a company building general intelligence for the physical world. GEN-family models process video, sensor data, proprioceptive data and language, and output robot actions (motion trajectories). The family includes GEN-1 (support for multiple end effectors and transfer across robotic interfaces) and GEN-1.5 (one-shot learning from a single demonstration and few-shot adaptation). The models are trained via scaled pretraining on large physical-interaction datasets collected in homes, warehouses and factories.

