Robots Atlas>ROBOTS ATLAS
K2 Horizon MoVA 36B-A4B

K2 Horizon MoVA 36B-A4B

MoVA 36B-A4Bย ยทย Family: K2 Horizon
Sparse MoE member of the K2 Horizon family (IFM): 36B parameters, 4B active per token, Mixture-of-Values attention, 512K context.
โœ“ Activeโœ“ Public accessโš– Open sourceLLMReasoning modelTool-using model๐Ÿ“ K2 Horizon
Context window
512K
tokens
Parameters
36B (4B aktywnych)
parameters
Access:DownloadDeployment:๐Ÿ’ป Localโ˜ Cloud

Overview

K2 Horizon MoVA 36B-A4B is the sparse member of the K2 Horizon model family from the Institute of Foundation Models (IFM). It is a Mixture-of-Experts (MoE) model with Mixture-of-Values attention (MoVA): it stores 36B parameters and activates about 4B per token.

The model provides a native 524,288-token context window (512K) from the midtraining stage onward. It is text-in / text-out, supports reasoning controlled via a reasoning-effort parameter, and tool use in agentic scenarios. Weights are released under Apache 2.0, and the training data, recipe and training code are stated to be released.

It is served via vLLM and SGLang in BF16 precision, with GGUF quantizations available (llama.cpp, Ollama, LM Studio, Jan). Reported scores were obtained at high reasoning effort.

Classification
LLMReasoning modelTool-using model
Family: K2 Horizon
Access & deployment
Download
LocalCloud
Weights: Open source
Key parameters
๐Ÿ“ Context: 512K
๐Ÿงฉ Parameters: 36B (4B aktywnych)
โœ“ Tools
๐Ÿ“ฅ Input: text

Technical specification

Context window
512K
tokens
Parameters
36B (4B aktywnych)
parameters
License
Apache 2.0
Hardware requirements
BF16; serving validated on 2x NVIDIA H200 (tensor-parallel=2) via vLLM/SGLang.
Features:โœ“ Tool use
Modalities
โฌ‡ Input
text
โฌ† Output
textcode

Capabilities and applications

Native model capabilities
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Advanced reasoning
The ability to perform multi-step, structured reasoning: analysing problems, planning steps, and drawing conclusions from hypotheses. Reasoning-first models (e.g. GPT-5.1 Thinking) dedicate a portion of inference to chains of thought before responding.
Category: reasoning
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Agentic coding
Multi-hour, multi-step programming tasks performed autonomously by the model: cloning a repository, running tests, iterating on fixes, integrating with CLI tools. Characteristic of Codex variants (GPT-5.1-Codex-Mini, Codex-Max).
Category: coding
Tool use
The model's ability to call external functions, APIs and tools during a conversation: calculator, search engine, code editor, database. The model decides when and how to use a tool and interprets its result.
Category: planning
Agentic capability
The model's ability to autonomously plan and execute multi-step tasks by sequentially using tools, maintaining context, and adapting to intermediate results.
Category: planning
Long context
Support for large context windows โ€” tens to hundreds of thousands (or millions) of input tokens. Enables analysis of entire codebases, long documents, and many parallel conversations without losing earlier information. GPT-5.1 supports 400,000 tokens.
Category: language

Benchmark results

9 benchmarks
tau3-Banking
high reasoning effort
26.8%
๐Ÿ“„ IFM
Agentic tool use
Terminal-Bench 2.0
high reasoning effort
58.6%
๐Ÿ“„ IFM
Agentic terminal use
SciCode
high reasoning effort
38.9%
๐Ÿ“„ IFM
Scientific coding
Humanity's Last Exam (HLE)
high reasoning effort
25.2%
๐Ÿ“„ IFM
GPQA
high reasoning effort
80.8%
๐Ÿ“„ IFM
CritPt
high reasoning effort
2.1%
๐Ÿ“„ IFM
Frontier physics reasoning
AA-LCR
high reasoning effort
66.3%
๐Ÿ“„ IFM
Long-context reasoning
AA-Omniscience Accuracy
18.8%
๐Ÿ“„ IFM
Factual accuracy
AA-Omniscience Non-Hallucination
69.2%
๐Ÿ“„ IFM
Non-hallucination rate

Technical architecture