Fourth-generation DeepSeek language models (V4-Flash and V4-Pro variants). MoE architecture, 1M-token context window. Preview: April 2026.
Context window
1M
tokens
Parameters
284B (V4-Flash) โ 1,6T (V4-Pro)
parameters
Release date
24 April 2026
Access:APIDownloadHostedDeployment:โ Cloud๐ป Local
Overview
Access & deployment
APIDownloadHosted
CloudLocal
Key parameters
๐ Context: 1M
๐งฉ Parameters: 284B (V4-Flash) โ 1,6T (V4-Pro)
Technical specification
Context window
1M
tokens
Parameters
284B (V4-Flash) โ 1,6T (V4-Pro)
parameters
License
MIT (V4-Flash) / proprietary (V4-Pro)
Capabilities and applications
Native model capabilities
Advanced reasoning
The ability to perform multi-step, structured reasoning: analysing problems, planning steps, and drawing conclusions from hypotheses. Reasoning-first models (e.g. GPT-5.1 Thinking) dedicate a portion of inference to chains of thought before responding.
Category: reasoning
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Long context
Support for large context windows โ tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Mathematical reasoning
The model's ability to solve mathematical tasks requiring multi-step reasoning โ equations, proofs, combinatorics, geometry, calculus and competition-level problems.
Category: reasoning
Technical architecture
Core Architecture
Sources and related pages
2 sources
Browse related topics
