Robots Atlas>ROBOTS ATLAS
Qwen3-VL

Qwen3-VL

Family: Qwen
Qwen's open vision-language model family from Alibaba: image, video and document understanding, OCR, grounding and a visual agent. Variants from 2B to 235B (dense and MoE).
✓ Active✓ Public access⚖ Open sourceMultimodalReasoning modelTool-using model📁 Qwen
Context window
256K (rozszerzalne do 1M przez YaRN)
tokens
Parameters
2B–235B (dense i MoE)
parameters
Release date
23 September 2025
Access:DownloadAPIHostedDeployment:💻 Local☁ Cloud

Overview

Qwen3-VL is a family of open vision-language models (VLM) developed by the Qwen team (Alibaba). The models accept text, images, video and documents as input and generate text, code and structured data as output.

The family includes dense variants at 2B, 4B, 8B and 32B, plus Mixture-of-Experts (MoE) variants at 30B-A3B and 235B-A22B. Each size ships in an Instruct (instruction-tuned) edition and a Thinking (reasoning-enhanced) edition, as well as FP8 quantization.

The native context window is 256K tokens and can be extended to 1M tokens via the YaRN technique. The model supports OCR in 32 languages, visual grounding (2D/3D boxes and points), long-video understanding with temporal grounding, document parsing, and operation as a visual agent controlling PC and mobile GUIs.

Architectural features include Interleaved-MRoPE (positional embeddings over time, width and height), DeepStack (fusing multi-level ViT features), and text–timestamp alignment. The models are released under the Apache-2.0 license.

Classification
MultimodalReasoning modelTool-using model
Family: Qwen
Access & deployment
DownloadAPIHosted
LocalCloud
Weights: Open source
Key parameters
📏 Context: 256K (rozszerzalne do 1M przez YaRN)
🧩 Parameters: 2B–235B (dense i MoE)
✓ Tools · ✓ Fine-tuning
📥 Input: text, image, video, documents

Technical specification

Context window
256K (rozszerzalne do 1M przez YaRN)
tokens
Parameters
2B–235B (dense i MoE)
parameters
License
Apache-2.0
Features:✓ Tool use✓ Fine-tuning
Modalities
⬇ Input
textimagevideodocuments
⬆ Output
textcodestructured_data

Capabilities and applications

Native model capabilities
Multimodal understanding
Category: multimodal
Image understanding
Analysing and interpreting the content of images.
Category: vision
Video understanding
The model's ability to analyse and interpret video content — recognising actions, motion, events and relationships between objects over time.
Category: video
OCR
Recognising text within images and documents.
Category: vision
Visual grounding
Locating and pointing to objects in an image in response to a text query — returning bounding boxes or points grounded to the description.
Category: vision
Chart understanding
Reading and interpreting charts, tables and diagrams.
Category: vision
Object tracking (video)
The ability to track selected objects across consecutive video frames, maintaining their masks/identity despite motion, occlusion and appearance changes.
Category: vision
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Computer use
The model's ability to operate a computer interface by interpreting screenshots and generating actions such as clicks, typing, and navigating applications.
Category: planning
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language

Pricing

Technical architecture

Core Architecture