Robots Atlas>ROBOTS ATLAS
QwenVL3.5-4B

QwenVL3.5-4B

QwenVL3.5-4B (wariant Qwen3.5-VL ~4B) · Family: Qwen
A 4B-class vision-language model from Alibaba's Qwen3.5-VL family, known primarily as the planning-layer backbone of PsiBot's Psi-R2.5 foundation model.
✓ Active⏳ Limited access⚖ Open weightsVisionMultimodalLLM📁 Qwen
Access:DownloadDeployment:💻 Local☁ Cloud

Overview

QwenVL3.5-4B is a 4-billion-parameter-class vision-language model from the Qwen3.5-VL family developed by Alibaba. Under this name it appears primarily in the technical documentation of the Chinese company PsiBot, which used it as the upper-layer (planning) backbone of its Psi-R2.5 embodied AI foundation model, unveiled in September 2026.

In that role the model processes vision-language instructions and decomposes them into subtasks, drawing on metadata context — memory and soft prompts — and on value signals from reinforcement learning. Its output feeds the lower layer built on Wan2.2-TI2V-5B, which turns the plan into robot action trajectories. PsiBot states that both layers leverage its own proprietary pre-training data.

A note on naming: the Qwen3.5-VL family does exist and is published by Alibaba in 2026 in several sizes (including 2B, 9B and 122B-A10B checkpoints, with the 9B version serving as the basis for Qwen-Image-2.1). Alibaba has not, however, published an official model card under the exact name "QwenVL3.5-4B" — that spelling comes from PsiBot's materials, which remain the only direct source for the 4B variant used in Psi-R2.5. The exact parameter count, context window and licence of this particular checkpoint have not been disclosed publicly.

Classification
VisionMultimodalLLM
Family: Qwen
Access & deployment
Download
LocalCloud
Weights: Open weights
Key parameters
📥 Input: text, image, video

Technical specification

Modalities
⬇ Input
textimagevideo
⬆ Output
textstructured_data

Capabilities and applications

Native model capabilities
Multimodal understanding
Category: multimodal
Image understanding
Analysing and interpreting the content of images.
Category: vision
Video understanding
The model's ability to analyse and interpret video content — recognising actions, motion, events and relationships between objects over time.
Category: video
Planning
Forming and executing action plans for complex tasks.
Category: planning
Instruction following
Precisely following instructions contained in the prompt: response format, length, style, constraints (e.g. 'reply in six words'). GPT-5.1 significantly improved this capability compared to GPT-5.
Category: language
Vision encoder
The model's ability to encode images and video frames into dense representations (embeddings), used for downstream tasks or as a backbone for vision-language models.
Category: vision