Experimental multimodal (image+text-to-text) model from DeepSeek's V4 family, built on DeepSeek-V4-Flash with added vision modules. Open weights, MIT license.
Context window
1M
tokens
Access:DownloadDeployment:๐ป Localโ Cloud
Overview
Access & deployment
Download
LocalCloud
Weights: Open source
Key parameters
๐ Context: 1M
โ Tools
๐ฅ Input: image, text
Technical specification
Context window
1M
tokens
License
MIT
Hardware requirements
Served via vLLM or SGLang; example config: a single 4รGB300 node (tensor-parallel-size 4), FP8/FP4 quantization.
Features:โ Tool use
Modalities
โฌ Input
imagetext
โฌ Output
textcode
Capabilities and applications
Native model capabilities
Image understanding
Analysing and interpreting the content of images.
Category: vision
Multimodal understanding
Category: multimodal
Chart understanding
Reading and interpreting charts, tables and diagrams.
Category: vision
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Agentic coding
Multi-hour, multi-step programming tasks performed autonomously by the model: cloning a repository, running tests, iterating on fixes, integrating with CLI tools. Characteristic of Codex variants (GPT-5.1-Codex-Mini, Codex-Max).
Category: coding
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Long context
Support for large context windows โ tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Vision encoder
The model's ability to encode images and video frames into dense representations (embeddings), used for downstream tasks or as a backbone for vision-language models.
Category: vision
Benchmark results
11 benchmarks
Terminal Bench 2.1
DeepSeek Harness minimal mode, max reasoning effort, temperature=1.0, top_p=0.95
83.9
๐ technical_report
NL2Repo
57.7
๐ technical_report
Cybergym
75.3
๐ technical_report
DeepSWE
59.3
๐ technical_report
Toolathlon-Verified
75.9
๐ technical_report
DSBench-Hard
63.6
๐ technical_report
AutomationBench (Public)
25.7
๐ technical_report
ApexBench
Pass@1
36.5
๐ technical_report
Agents' Last Exam
27.3
๐ technical_report
Chartography
64.3
๐ technical_report
ZeroBench
Pass@5
35.0
๐ technical_report
Technical architecture
Core Architecture
Model Form
Training Techniques
