Quick Facts
- Google Research and Virginia Tech published the WikiSkill framework on arXiv on Aug. 27, 2026, giving AI agents a structured knowledge layer that persists between task runs.
- Gemini 3.5 Flash's average score across five benchmarks jumped from 49.5% to 68.1% with WikiSkill applied.
- Skills evolved by smaller models can transfer to larger ones, with Qwen-3.5-4B skills pushing Gemma-4-31B to 73.1% on the LiveMath benchmark.
AI agents built on large language models forget everything when a task ends. Run the same agent twice on a similar problem and it often repeats the same mistake. WikiSkill, a new framework from Google Research and Virginia Tech, is designed to stop that cycle.
Authored by Liyan Tang, Cyrus Rashtchian, and four colleagues, the paper was published on arXiv on Aug. 27, 2026. It introduces a three-layer architecture that captures what an agent did, distills what worked and what failed, and packages that knowledge into reusable modules called Agent Skills.
How the System Works
The raw layer stores complete execution traces, including every tool call and its result. That data is immutable. Above it, the wiki layer distills those traces into structured notes on failure patterns and successful strategies. The wiki never resets; it only grows.
The skill layer sits on top. Instead of loading lessons into the prompt, the system packages validated procedures into Agent Skills, reusable modules that guide future behavior without altering the model's weights or bloating its context window.
The system runs on a four-step loop. An inference agent attempts tasks using current skills. A wiki maintainer analyzes successful and failed runs and updates the knowledge base. A skill proposer reads the updated wiki and suggests one change to an existing skill or drafts a new one. That change only sticks if it beats the previous best score on a separate set of validation tasks.
A Guard Against Skill Degradation
Systems that let agents rewrite their own procedures can degrade fast if every revision is accepted. WikiSkill's validation gate is built to prevent that. If a proposed skill fails to improve validation performance, the system discards the skill update but keeps the underlying wiki notes.
One concrete example from the research: an early proposed skill was rejected for being too abstract. A later accepted rule reads, "Never Return an Item to Its Origin Location." The specificity matters. Vague instructions do not pass the gate.
Benchmark Results
The team tested WikiSkill across five benchmarks covering math reasoning, web search, spreadsheet manipulation, document question-answering, and interactive virtual environments. They ran it against five models.
Gemini 3.5 Flash improved from 49.5% to 68.1% on average. On the LiveMath benchmark alone, the same model jumped from 33.0% to 72.6%. On SpreadSheet tasks, it rose from 50.5% to 76.6%. Qwen-3.6-27B moved from 39.4% to 63.3% on average. Gains scaled with model size: 12.3, 17.5, and 23.9 points for Qwen's 4B, 9B, and 27B variants, respectively.
WikiSkill outperformed three existing skill-updating tools and standard no-skill setups across most settings.
Cross-Model Skill Transfer
One of the more unexpected findings involves skill portability. Skills developed by one model can transfer to another, and smaller models can produce skills that improve larger ones. Qwen-3.5-4B skills pushed Gemma-4-31B to 73.1% on LiveMath and 66.9% on ALFWorld. The researchers concluded that stronger source models do not necessarily produce better transferable skills.
Lead researcher Liyan Tang described WikiSkill's adaptation of Andrej Karpathy's LLM Wiki concept, which Karpathy published in early 2026, this way: "We kept Karpathy's shape: immutable sources, an LLM-maintained wiki, an index and log, but the input is the agent's own execution trajectory, and the output is an executable SKILL.md."
What It Means for Enterprise AI Teams
For teams deploying agents in production, WikiSkill offers a path to turning execution traces, data most organizations already generate, into reusable procedural knowledge. A coding assistant, research agent, or document processor could preserve validated procedures and document unsuccessful tool calls without retraining the underlying model.
The framework points to a practical alternative when an agent repeatedly fails on the same class of task. Rather than fine-tuning or prompt engineering, teams could let the wiki accumulate and let the validation gate filter what actually improves performance.
Read more: Google's WikiSkill gives AI agents a memory of what went wrong without putting it in the prompt
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
