Robots Atlas>ROBOTS ATLAS
K2 Horizon 375B-A23B

K2 Horizon 375B-A23B

375B-A23Bย ยทย Family: K2 Horizon
Flagship of IFM K2 Horizon family: a sparse MoE with 375B total parameters (23B active), 512K context, Apache 2.0, fully open.
โœ“ Activeโœ“ Public accessโš– Open sourceLLMReasoning model๐Ÿ“ K2 Horizon
Context window
512K
tokens
Parameters
375B-A23B (375 mld total / 23 mld aktywne, MoE)
parameters
Release date
3 September 2026
Access:DownloadDeployment:๐Ÿ’ป Localโ˜ Cloud

Overview

K2 Horizon 375B-A23B is the flagship and largest model of the K2 Horizon family developed by the Institute of Foundation Models (IFM), released on 3 September 2026. It uses a sparse Mixture-of-Experts (MoE) architecture: it stores 375B parameters and activates roughly 23B per token.

The model supports a native 512K context window (524,288 tokens) from the midtraining stages onward. It targets demanding workloads: complex reasoning, software engineering, research, and long-horizon agentic tasks. By default it runs with high reasoning effort.

K2 Horizon 375B-A23B is fully open following the LLM360 principle: beyond the final checkpoint, IFM plans to release intermediate checkpoints, data, the training recipe, and code. Weights are available under Apache 2.0 in safetensors (BF16) format, with an official FP8 quantized release. Serving was validated on a single 8ร— H200 node (vLLM, SGLang, Transformers).

Classification
LLMReasoning model
Family: K2 Horizon
Access & deployment
Download
LocalCloud
Weights: Open source
Key parameters
๐Ÿ“ Context: 512K
๐Ÿงฉ Parameters: 375B-A23B (375 mld total / 23 mld aktywne, MoE)
โœ“ Toolsย ยทย โœ“ Fine-tuning
๐Ÿ“ฅ Input: text

Technical specification

Context window
512K
tokens
Parameters
375B-A23B (375 mld total / 23 mld aktywne, MoE)
parameters
License
Apache 2.0
Hardware requirements
BF16 serving on a single 8ร— H200 node (tensor-parallel 8, expert-parallel), e.g. via vLLM or SGLang. An official FP8 release lowers memory requirements.
Features:โœ“ Tool useโœ“ Fine-tuning
Modalities
โฌ‡ Input
text
โฌ† Output
textcode

Capabilities and applications

Native model capabilities
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Advanced reasoning
The ability to perform multi-step, structured reasoning: analysing problems, planning steps, and drawing conclusions from hypotheses. Reasoning-first models (e.g. GPT-5.1 Thinking) dedicate a portion of inference to chains of thought before responding.
Category: reasoning
Mathematical reasoning
The model's ability to solve mathematical tasks requiring multi-step reasoning โ€” equations, proofs, combinatorics, geometry, calculus and competition-level problems.
Category: reasoning
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Agentic coding
Multi-hour, multi-step programming tasks performed autonomously by the model: cloning a repository, running tests, iterating on fixes, integrating with CLI tools. Characteristic of Codex variants (GPT-5.1-Codex-Mini, Codex-Max).
Category: coding
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Long context
Support for large context windows โ€” tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Adaptive reasoning effort
The model decides how much 'thinking' to allocate to a given query: simple questions are answered quickly, complex problems receive more inference cycles. A GPT-5.1 feature (both Instant and Thinking) that shortens time on easy tasks and extends it for hard ones.
Category: reasoning

Benchmark results

16 benchmarks
GPQA
accuracy ยท high reasoning effort
87.3%
๐Ÿ“„ IFM
Humanity's Last Exam (HLE)
accuracy ยท high reasoning effort
32.0%
๐Ÿ“„ IFM
Terminal-Bench 2.0
accuracy ยท high reasoning effort
70.2%
๐Ÿ“„ IFM
SWE-Bench Pro
accuracy ยท high reasoning effort
42.6%
๐Ÿ“„ IFM
GDPval-AA
Elo ยท high reasoning effort
1441Elo
๐Ÿ“„ IFM
SciCode
accuracy ยท high reasoning effort
42.7%
๐Ÿ“„ IFM
MCPMark
accuracy ยท high reasoning effort
67.7%
๐Ÿ“„ IFM
Toolathlon Verified
accuracy ยท high reasoning effort
65.3%
๐Ÿ“„ IFM
BrowseComp
accuracy ยท high reasoning effort
72.8%
๐Ÿ“„ IFM
tau3-Banking
accuracy ยท high reasoning effort
34.0%
๐Ÿ“„ IFM
Apex-Agents (pass@1)
accuracy ยท high reasoning effort
24.8%
๐Ÿ“„ IFM
Automation Bench Public
accuracy ยท high reasoning effort
25.3%
๐Ÿ“„ IFM
WildClawBench
accuracy ยท high reasoning effort
50.9%
๐Ÿ“„ IFM
CritPt
accuracy ยท high reasoning effort
8.6%
๐Ÿ“„ IFM
AA-LCR
accuracy ยท high reasoning effort
76.0%
๐Ÿ“„ IFM
SWE-Atlas-QnA (strict)
accuracy ยท high reasoning effort
48.4%
๐Ÿ“„ IFM

Technical architecture