Robots Atlas>ROBOTS ATLAS
Artificial Intelligence

Raven 0.2.0: EverMind Orchestrates Claude Code and Codex

Sir Robot2 October 2026 · 3 min read
Raven 0.2.0: EverMind Orchestrates Claude Code and Codex

EverMind AI released Raven 0.2.0 on 23 September, an open-source framework calling itself “the harness of harnesses”. Its claim: an agent improves without touching model weights, if it may rewrite its own control layer. An 82-page technical report followed four days later.

Key takeaways

  • Raven 0.2.0 shipped 23 September 2026 under Apache-2.0
  • The agent loop splits into four modules — Memory, Planning, Capability, Action — rewritten by a Curator
  • 13 third-party agent presets, among them Claude Code, Codex and GitHub Copilot
  • Raven RSI: 172 training runs across 7 rounds with no crash, val_bpb down 5.8%

Self-improvement moved into the control layer

Raven splits the agent loop into four strategy modules. Above them the Curator: Raven's experimental module that rewrites the agent's four strategy modules, either as a setting or as a small piece of judgement code. rewrites those four seats — sometimes a setting, sometimes a small piece of judgement code.

Strategy moduleWhat it covers
MemoryWhat a turn sees
PlanningHow the work is approached
CapabilityWhich tools are exposed
ActionPicking the next move and judging it before it runs

Two safeguards hold. Nothing reaches the agent unverified, and a round's signals arrive with the reference answer stripped out. The Curator is experimental and ships with the repository, not the package.

Thirteen agents under one conductor

The other half is orchestration. Raven acts as a Host Agent: it breaks a goal into subtasks, matches agents to them, tracks dependencies and merges results.

Four agents are its own — Raven-Research, Raven-Code, Raven-Design and Raven-Oncall. The rest connect over ACP: Agent Client Protocol — an open JSON-RPC protocol through which a host program or editor talks to an external coding agent. Not to be confused with the Agent Communication Protocol, which shares the acronym., CLI or OpenAI-compatible APIs. Thirteen presets ship, led by Claude Code and Codex.

Orchestration
Human goal → Host Agent
Agents matched: Claude Code, Codex, Raven-Code
Subtasks run, results merged
Self-improvement
Curator rewrites Memory, Planning, Capability, Action
Does the change pass verification?
YES
Installed in the agent's harnessAllow
NO
Rejected, back to the CuratorDeny
Self-improvement
Next round on the improved harnessAllow

Numbers from the home bench

Raven RSI is the strongest demonstration: handed a task and fixed criteria, the system planned its own rounds, wrote code and ran experiments. On nanochat pre-training it cut val_bpb by 5.8% across 172 runs without a single crash.

−5.8%val_bpb across 172 nanochat pre-training runs, with no crashEverMind AI

Salesforce tests the same thesis with DarwinX, which lifted an agent from 43.5% to 93% on WebArena-Infinity. The difference is scope: DarwinX polishes one agent's harness, Raven a team of outside tools. The figures are EverMind’s own. Raven is pre-alpha.

Why it matters

If progress comes from rewriting the control layer instead of another training job, the cost of entry into self-improving systems drops from a GPU farm to an ordinary server.

Raven pushes further than a single agent: it treats other vendors' tools as swappable parts of its own architecture. That casts model providers as subcontractors to a layer they do not control.

What's next

  • Moving the Curator out of the repository into the installed package would signal EverMind considers it stable
  • 0.2.1, 0.2.2 and 0.2.3 followed within four days, the last on 27 September
  • The report's claim that composing agents widens task coverage beyond the individual parts awaits independent checking

Sources

Share this article