Quick Facts
- SIFT, developed by MIT and Sakana AI, reduces self-improving coding agent search costs to as low as $34 in API spend, compared to roughly $22,000 for predecessor method Darwin Gödel Machine.
- On the Polyglot coding benchmark, SIFT reached 35.1% accuracy in under five hours of wall-clock time using about $150 in API credits.
- On a 60-task SWE-bench Verified subset, SIFT pushed a gpt-5-mini agent from 51.7% to 61.7% in fewer than three self-improvement steps at a cost of $25.
Researchers at MIT and Sakana AI have published a framework that cuts the cost of self-improving coding agents by an order of magnitude, opening a class of AI research that was previously accessible only to well-funded labs.
The framework, called Recursive Self-Improvement via Fast Tree Search, or SIFT, was authored by MIT's Xinghong Fu with collaborators Aravinth Kulanthaivelu and Yutaro Yamada at Sakana AI. The paper appeared on arXiv around September 18, 2026.
The Problem With Prior Methods
Self-improving agents work by modifying the prompts, tools, and code that guide their behavior. The bottleneck in earlier systems was determining which modifications actually help. Previous approaches re-ran subsets of benchmark tasks to measure each candidate change, a process that is slow and expensive.
The authors describe the core issue as "the poor signal-to-cost trade-off of existing evaluation methods." Evaluating one agent version on 50 Polyglot tasks cost about $6 and 2.6 CPU hours. Completing a single run of the Darwin Gödel Machine, the closest predecessor, cost an estimated $22,000 and took about two weeks.
How SIFT Works
SIFT replaces full benchmark re-runs with pairwise LLM-as-a-judge comparisons. After a self-improvement model writes a new agent version, a judge model compares its source files against up to 10 strong incumbents from the current archive. Each comparison returns a single preference signal at a cost of about 4.4 cents per call.
Those preferences are aggregated using a Bradley-Terry model, a statistical technique that converts noisy pairwise outcomes into a global ranking. SIFT uses only the rank order, not raw scores, making the system less sensitive to the number of comparisons accumulated by any individual candidate.
Candidate generation, judging, and more expensive evaluations run in parallel inside a disaggregated tree-search pipeline. That parallel structure is where most of SIFT's speed advantage comes from.
Benchmark Results
On Polyglot, a SIFT run using an o3-mini agent reached 35.1% accuracy in 30 expansion steps, compared to 30.7% for DGM over 80 nodes. The full run finished in under five hours using 42 CPU hours and about $150 in API credits.
On TerminalBench 2.1, scores climbed from 29.2% to 36.7%. On SWE-60, performance rose from 40.0% to 52.1%. A configuration built on Qwen3-Coder-30B finished its entire search in 224 CPU hours at approximately $34 in API costs, roughly one-tenth of DGM's resource consumption.
The researchers also report that SIFT produces agent harnesses that transfer across coding models, rather than ones tuned specifically for the model used during search.
What This Means for Software Teams
Sakana AI is a Tokyo-based lab co-founded by Llion Jones, one of the authors of the original "Attention Is All You Need" paper. The lab recently established a dedicated Recursive Self-Improvement Lab in Tokyo and named AI researcher Jurgen Schmidhuber as its chief scientific adviser.
The cost reduction SIFT demonstrates has a direct business implication. Self-improving agent research has been the province of labs with large compute budgets. At $25 to $150 per run, the technique becomes viable for smaller engineering teams evaluating whether automated agent improvement fits their development pipelines.
For founders and CTOs watching AI development costs, SIFT signals that the price of building agents that improve themselves is falling fast and may reach mainstream viability sooner than expected.
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
