Robots Atlas>ROBOTS ATLAS
OpenAegis

OpenAegis

OpenAegis (Qwen 3.5-397B-A17B SFT)
Open cybersecurity LLM (SFT on Qwen 3.5-397B-A17B) trained with the CyberFactory framework; 58.1% Pass@1 on CyberGym.
๐Ÿ”ฌ Research๐Ÿ”ฌ Research onlyโš– Open weightsLLMSpecialized AITool-using model
Parameters
397B total / 17B active (MoE)
parameters
Release date
1 September 2026
Access:DownloadDeployment:๐Ÿ’ป Localโ˜ Cloud

Overview

OpenAegis is a cybersecurity-specialized large language model introduced in the paper "CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild" (arXiv:2608.23181). It is produced via full-parameter supervised fine-tuning (SFT, three epochs) of the Qwen 3.5-397B-A17B base model (a mixture-of-experts model with 397B total and 17B active parameters).

The model is trained on trajectories generated by CyberFactory, an open-source framework that connects data construction, trajectory synthesis, and model training across three tasks: proof-of-concept (PoC) generation, vulnerability patching, and cybersecurity question answering (CyberQA). On the CyberGym benchmark, OpenAegis reaches 58.1% Pass@1 under a one-hour per-task budget, improving over its Qwen 3.5 base by 28.5 points.

Classification
LLMSpecialized AITool-using model
Access & deployment
Download
LocalCloud
Weights: Open weights
Key parameters
๐Ÿงฉ Parameters: 397B total / 17B active (MoE)
โœ“ Tools
๐Ÿ“ฅ Input: text

Technical specification

Parameters
397B total / 17B active (MoE)
parameters
Features:โœ“ Tool use
Modalities
โฌ‡ Input
text
โฌ† Output
textcode

Capabilities and applications

Native model capabilities
Coding
Generating, analysing and modifying code in many programming languages. Covers writing functions, debugging, refactoring, code review, and creating tests. Measured by benchmarks such as HumanEval and SWE-bench.
Category: coding
Reasoning
The model's ability to reason logically and solve complex problems.
Category: reasoning
Cybersecurity
The model ability to perform computer-security tasks: vulnerability analysis, proof-of-concept exploit generation, patching, and cybersecurity question answering.
Category: other
Vulnerability detection
The ability to identify and analyze security vulnerabilities in source code and software.
Category: coding

Benchmark results

1 benchmark
CyberGym
Pass@1 ยท One-hour per-task budget; +28.5 points over the Qwen 3.5 base model.
58.1%
๐Ÿ“… 1 Sept 2026๐Ÿ“„ paper

Technical architecture

Core Architecture
Model Form
Training Techniques