Tencent MoE model (770B parameters, 49B active) with a 1M-token context window, optimized for software engineering, office productivity and research.
Context window
1M
tokens
Parameters
770B (49B aktywnych)
parameters
Access:DownloadDeployment:☁ Cloud💻 Local
Overview
Applications
Access & deployment
Download
CloudLocal
Weights: Open weights
Key parameters
📏 Context: 1M
🧩 Parameters: 770B (49B aktywnych)
✓ Tools · ✓ Fine-tuning
📥 Input: text
Technical specification
Context window
1M
tokens
Parameters
770B (49B aktywnych)
parameters
License
Apache 2.0
Hardware requirements
The 770B MoE model (49B active) requires a multi-GPU data-center cluster (e.g. multiple NVIDIA H100/H200) for full-precision inference; quantization lowers the requirements.
Features:✓ Tool use✓ Fine-tuning
Modalities
⬇ Input
text
⬆ Output
textcode
Capabilities and applications
Native model capabilities
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Agentic coding
Multi-hour, multi-step programming tasks performed autonomously by the model: cloning a repository, running tests, iterating on fixes, integrating with CLI tools. Characteristic of Codex variants (GPT-5.1-Codex-Mini, Codex-Max).
Category: coding
Mathematical reasoning
The model's ability to solve mathematical tasks requiring multi-step reasoning — equations, proofs, combinatorics, geometry, calculus and competition-level problems.
Category: reasoning
Advanced reasoning
The ability to perform multi-step, structured reasoning: analysing problems, planning steps, and drawing conclusions from hypotheses. Reasoning-first models (e.g. GPT-5.1 Thinking) dedicate a portion of inference to chains of thought before responding.
Category: reasoning
Long context
Support for large context windows — tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language
Spreadsheet generation
Generating complete spreadsheets (Excel, Google Sheets, CSV) with formulas, conditional formatting, pivot tables, and multi-sheet data structures based on a natural-language description.
Category: structured_generation
Presentation creation
Generating presentations (PowerPoint, Google Slides, Keynote): slide structure, text content, graphic suggestions, and a talking-points outline. The model picks a layout and narrative appropriate to the presentation's purpose.
Category: structured_generation
Benchmark results
3 benchmarks
GPQA
accuracy · GPQA Diamond
92.3%
📄 Oficjalna karta modelu (Hugging Face)
Terminal-Bench 2.0
accuracy · Terminal-Bench 2.1
85.4%
📄 Oficjalna karta modelu (Hugging Face)
SWE-bench
resolved · SWE-bench Multilingual (Resolved)
82.9%
📄 Oficjalna karta modelu (Hugging Face)
Technical architecture
Core Architecture
Model Form
