Happy Sunday. The price of self-improving AI agents just dropped by a factor of 650. A new framework from MIT and Sakana AI brings the cost of a full self-improvement run to as little as $34, down from $22,000 for the prior state of the art. That puts a capability once reserved for large research labs within reach of a small engineering team.

Alongside that, Amazon dropped a free, open source decision model purpose-built for agent pipelines, and Kevin Mandia's AI security startup Armadin closed a $255.5 million Series B at a $2.5 billion valuation, seven months after coming out of stealth. The thread across all three: the infrastructure for autonomous agents is being built fast, from the model layer to the security layer.


ARTIFICIAL INTELLIGENCE

Self-Improving Coding Agents Now Cost as Little as $34 to Run

MIT and Sakana AI Cut Self-Improving Agent Costs From $22,000 to $34

A new framework called SIFT, published by MIT and Sakana AI around Sept. 18, replaces expensive benchmark re-runs with cheap pairwise judge comparisons to evaluate agent improvements. On the SWE-bench Verified subset, it pushed a gpt-5-mini agent from 51.7% to 61.7% accuracy at a cost of $25. The previous leading method, Darwin Gödel Machine, cost an estimated $22,000 per run and took two weeks.

The practical ceiling for this kind of work was not algorithmic. It was compute budget. At $25 to $150 per run, teams can test whether automated agent improvement fits their pipelines without a large research budget. One result worth watching: SIFT produces agent harnesses that transfer across coding models, not ones tuned to a specific model. That means a single search run could benefit a whole stack rather than one deployment.

Read the story →


ARTIFICIAL INTELLIGENCE

Amazon's Free 2B Decision Model Is Built to Gate Your Agent Pipelines

Amazon Releases Free Open Source Decision Model That Runs in Under 100 Milliseconds

Amazon released Strands Decider 2B on Oct. 1 under an Apache-2.0 license. The 1.9-billion-parameter model does not generate text. It takes a list of options and returns a selection with a calibrated confidence score in a single forward pass, with median latency of 106 milliseconds on an NVIDIA RTX 3090. On JevBench, it ranks first among public models that ship with a full training recipe.

The architectural argument here matters more than the benchmark number. AWS is proposing a two-tier agent design: a small, fast model handles routing and gating, while full generative calls go to a large model only when needed. That split could cut costs sharply in pipelines where most decisions are binary. One production note to flag: the bundled HTTP server ships with no authentication and binds to 127.0.0.1, so teams cannot use it as-is in production without adding their own security layer.

Read the story →


SECURITY & PRIVACY

Armadin Hits $2.5B Valuation Seven Months After Launch

Kevin Mandia's AI Security Startup Armadin Raises $255.5M at $2.5B Valuation

Armadin, the autonomous offensive security company co-founded by former Mandiant CEO Kevin Mandia, closed a $255.5 million Series B on Oct. 1, co-led by Andreessen Horowitz and Accel. The round values the company at over $2.5 billion. In a three-day live exercise with no privileged access granted at the start, Armadin Red agents produced 238 security findings and validated 38 complete attack chains.

The funding pace is the signal. Armadin raised $445 million total in under a year, a figure that surpasses reported totals for competitors Horizon3 and XBOW. The bet is that annual penetration tests are structurally unable to keep pace with attacks running at machine speed. For CISOs at software companies, the question is not whether to move toward continuous automated testing but how fast that shift becomes the expectation rather than the edge case.

Read the story →

Created by the robots at The SaaS Sentinel