Robots Atlas>ROBOTS ATLAS
Artificial Intelligence

WikiSkill: agents log their own failures in a wiki, not the prompt

Sir Robot10 October 2026 · 3 min read
WikiSkill: agents log their own failures in a wiki, not the prompt

Google Research and Virginia Tech have described WikiSkill, a framework that keeps an agent's lessons from failed attempts in a persistent knowledge base rather than the prompt. The paper went up on arXiv on 27 August 2026, with a broader analysis from VentureBeat on 1 October. On the largest model tested, the gain reached 23.9 points of average accuracy.

Key takeaways

  • Three layers: Raw (immutable execution traces), Wiki (consolidated knowledge), Skill (executable instructions)
  • Five benchmarks: LiveMath, SealQA, SpreadSheet, OfficeQA and ALFWorld
  • Five models: Qwen-3.5-4B, Qwen-3.5-9B, Qwen-3.6-27B, Gemma-4-31B and Gemini 3.5 Flash
  • Margin over the strongest competing method: 3.3 to 12.0 points of average accuracy
  • Qwen-3.5-9B with WikiSkill hits 47.4%, Qwen-3.6-27B without skills manages 39.4%
+23.9 pointsaverage accuracy on Qwen-3.6-27B versus running with no skillsarXiv 2608.27454

The wiki stays out of the prompt

Earlier Skill evolution: Automatically discovering reusable procedures from an agent's execution log and refining them across successive runs. methods lost the reason for a failure: the insight was scattered across optimisation history, so the agent hit the same wall repeatedly.

Raw
Record the task execution
Wiki
Consolidate into the wiki
Skill
Propose a new skill revision
Does the candidate beat validation?
YES
New active skill setAllow
NO
Candidate rejectedDeny
Next iteration on a new task

The cost split is the point: the wiki can grow without bound yet never enters the context on every call — the agent gets only the procedural digest. Google models produce terser skills (45.1 lines for Gemma-4-31B, 81.2 for Gemini 3.5 Flash) than the Qwen models, at 118.9 to 128.6 lines.

Bigger model, bigger payoff

Inside the Qwen family the improvement scales with size. Skills can also substitute for parameters — Qwen-3.5-9B with WikiSkill averages 47.4% and beats the three-times-larger Qwen3.6-27B without skills, which stops at 39.4%.

Modelvs. strongest competing methodvs. no skills
Qwen-3.5-4B+3.3 pts+12.3 pts
Qwen-3.5-9B+5.1 pts+17.5 pts
Qwen-3.6-27B+10.0 pts+23.9 pts
Gemma-4-31B+5.8 pts—
Gemini 3.5 Flash+12.0 pts—

Transfer can be toxic

Skills move between model families, sometimes beating home-grown ones — Qwen-3.5-9B reaches 70.2% on ALFWorld with instructions evolved by the 27B model, against 63.4% with its own.

Negative transfer. Skills from Qwen-3.5-4B dropped Gemini 3.5 Flash on SpreadSheet from 50.5% to 18.1% — they encoded workarounds built for a weaker model and burned through the tool-call budget. The same skills taken from the 27B model lifted that score to 63.4%.

Why it matters

Most agent deployments solve the memory problem through context engineering — stuffing the window, which costs tokens and dilutes the model's attention. Separating the archive from the executable instructions shifts the weight from inference to an offline consolidation step. The transfer results also show that skills are a swappable artefact, and that they can do harm when they come from a weaker model.

What's next

  • The paper gives no date for releasing the code or the wiki as a usable tool
  • Negative transfer on SpreadSheet is unresolved and still requires hand-picking the skill source
  • Gemini 3.5 Flash scored 100% on the ALFWorld validation split before evolution, leaving no baseline there

Sources

Share this article