Google Research and Virginia Tech have described WikiSkill, a framework that keeps an agent's lessons from failed attempts in a persistent knowledge base rather than the prompt. The paper went up on arXiv on 27 August 2026, with a broader analysis from VentureBeat on 1 October. On the largest model tested, the gain reached 23.9 points of average accuracy.
Key takeaways
- Three layers: Raw (immutable execution traces), Wiki (consolidated knowledge), Skill (executable instructions)
- Five benchmarks: LiveMath, SealQA, SpreadSheet, OfficeQA and ALFWorld
- Five models: Qwen-3.5-4B, Qwen-3.5-9B, Qwen-3.6-27B, Gemma-4-31B and Gemini 3.5 Flash
- Margin over the strongest competing method: 3.3 to 12.0 points of average accuracy
- Qwen-3.5-9B with WikiSkill hits 47.4%, Qwen-3.6-27B without skills manages 39.4%
The wiki stays out of the prompt
Earlier skill-evolution?Skill evolution: Automatically discovering reusable procedures from an agent's execution log and refining them across successive runs. methods lost the reason for a failure: the insight was scattered across optimisation history, so the agent hit the same wall repeatedly.
The cost split is the point: the wiki can grow without bound yet never enters the context on every call — the agent gets only the procedural digest. Google models produce terser skills (45.1 lines for Gemma-4-31B, 81.2 for Gemini 3.5 Flash) than the Qwen models, at 118.9 to 128.6 lines.
Bigger model, bigger payoff
Inside the Qwen family the improvement scales with size. Skills can also substitute for parameters — Qwen-3.5-9B with WikiSkill averages 47.4% and beats the three-times-larger Qwen3.6-27B without skills, which stops at 39.4%.
| Model | vs. strongest competing method | vs. no skills |
|---|---|---|
| Qwen-3.5-4B | +3.3 pts | +12.3 pts |
| Qwen-3.5-9B | +5.1 pts | +17.5 pts |
| Qwen-3.6-27B | +10.0 pts | +23.9 pts |
| Gemma-4-31B | +5.8 pts | — |
| Gemini 3.5 Flash | +12.0 pts | — |
Transfer can be toxic
Skills move between model families, sometimes beating home-grown ones — Qwen-3.5-9B reaches 70.2% on ALFWorld with instructions evolved by the 27B model, against 63.4% with its own.
Why it matters
Most agent deployments solve the memory problem through context engineering — stuffing the window, which costs tokens and dilutes the model's attention. Separating the archive from the executable instructions shifts the weight from inference to an offline consolidation step. The transfer results also show that skills are a swappable artefact, and that they can do harm when they come from a weaker model.
What's next
- The paper gives no date for releasing the code or the wiki as a usable tool
- Negative transfer on SpreadSheet is unresolved and still requires hand-picking the skill source
- Gemini 3.5 Flash scored 100% on the ALFWorld validation split before evolution, leaving no baseline there





