EverMind AI released Raven 0.2.0 on 23 September, an open-source framework calling itself “the harness of harnesses”. Its claim: an agent improves without touching model weights, if it may rewrite its own control layer. An 82-page technical report followed four days later.
Key takeaways
- Raven 0.2.0 shipped 23 September 2026 under Apache-2.0
- The agent loop splits into four modules — Memory, Planning, Capability, Action — rewritten by a Curator
- 13 third-party agent presets, among them Claude Code, Codex and GitHub Copilot
- Raven RSI: 172 training runs across 7 rounds with no crash, val_bpb down 5.8%
Self-improvement moved into the control layer
Raven splits the agent loop into four strategy modules. Above them the Curator?Curator: Raven's experimental module that rewrites the agent's four strategy modules, either as a setting or as a small piece of judgement code. rewrites those four seats — sometimes a setting, sometimes a small piece of judgement code.
| Strategy module | What it covers |
|---|---|
| Memory | What a turn sees |
| Planning | How the work is approached |
| Capability | Which tools are exposed |
| Action | Picking the next move and judging it before it runs |
Two safeguards hold. Nothing reaches the agent unverified, and a round's signals arrive with the reference answer stripped out. The Curator is experimental and ships with the repository, not the package.
Thirteen agents under one conductor
The other half is orchestration. Raven acts as a Host Agent: it breaks a goal into subtasks, matches agents to them, tracks dependencies and merges results.
Four agents are its own — Raven-Research, Raven-Code, Raven-Design and Raven-Oncall. The rest connect over ACP?ACP: Agent Client Protocol — an open JSON-RPC protocol through which a host program or editor talks to an external coding agent. Not to be confused with the Agent Communication Protocol, which shares the acronym., CLI or OpenAI-compatible APIs. Thirteen presets ship, led by Claude Code and Codex.
Numbers from the home bench
Raven RSI is the strongest demonstration: handed a task and fixed criteria, the system planned its own rounds, wrote code and ran experiments. On nanochat pre-training it cut val_bpb by 5.8% across 172 runs without a single crash.
Salesforce tests the same thesis with DarwinX, which lifted an agent from 43.5% to 93% on WebArena-Infinity. The difference is scope: DarwinX polishes one agent's harness, Raven a team of outside tools. The figures are EverMind’s own. Raven is pre-alpha.
Why it matters
If progress comes from rewriting the control layer instead of another training job, the cost of entry into self-improving systems drops from a GPU farm to an ordinary server.
Raven pushes further than a single agent: it treats other vendors' tools as swappable parts of its own architecture. That casts model providers as subcontractors to a layer they do not control.
What's next
- Moving the Curator out of the repository into the installed package would signal EverMind considers it stable
- 0.2.1, 0.2.2 and 0.2.3 followed within four days, the last on 27 September
- The report's claim that composing agents widens task coverage beyond the individual parts awaits independent checking
Sources
- GitHub — Raven 0.2.0
- GitHub — EverMind-AI/Raven
- GitHub — Raven Technical Report v1
- Jiqizhixin — Claude Code、Codex组队干活,Raven新版本既做总调度,也让Harness持续进化





