Quick Facts
- The share of enterprises allowing autonomous AI agents to make production changes without human review dropped from 75% in July to 56% in August 2026, per VB Intelligence.
- 61% of organizations running pre-deployment evaluations reported at least one AI agent that passed internal testing but caused a customer-facing failure in the past 12 months.
- Only 5% of enterprises say they fully trust automated evaluations, even as 66% already permit or are building toward unsupervised production deployment.
Enterprise confidence in autonomous AI agents is cracking. VB Intelligence's August Agentic Reliability and Evaluations survey found that 56% of respondents whose organizations deploy autonomous agents allow certain production changes without human review, or are engineering toward it within 12 months. That figure stood at 75% just one month earlier.
The shift is not a statistical artifact. Among final AI purchasing decision-makers specifically, the share fell from 88% in July to 61% in August. The share of enterprises planning to retain human review for the foreseeable future more than doubled, from 20% to 42%.
Failures Are Driving the Pullback
Real-world breakdowns are pushing organizations to slow down. Of the 140 August respondents, 61% whose companies run pre-deployment evaluations reported at least one case where an agent or LLM feature passed internal testing and then caused a customer-facing failure in the past year. A quarter of those companies saw it happen more than once.
The root problem is evaluation quality. Only 5% of enterprises say they fully trust automated evaluations. The most commonly cited limitation, flagged by 29% of respondents, is that evaluations align poorly with real-world outcomes. A passing test, enterprises are learning, does not guarantee a working agent.
The Gap Between Autonomy and Assurance
The data reveals a striking contradiction. Despite knowing their evaluations are unreliable, 66% of enterprises are already permitting some unsupervised production deployment or building toward it. Confidence in AI agents is outpacing the controls companies have built to govern them.
Harness, which released its own State of Agent DLC 2026 report alongside the VB findings, found that while 75% of organizations have AI agents running in production, actual controls around testing, security, cost, and rollback readiness lag confidence levels by 30 to 55 percentage points. Keith Mann, field CTO and head of research at Harness, called out the consistency problem directly. "A control that worked in testing can still miss something in production because an agent doesn't behave the same way every time," he said.
Trevor Stuart, SVP and general manager at Harness, said teams are now scrambling to govern what they have already shipped. "Teams moved fast to build and release agents, and are now circling back to ask how to actually govern what they've already shipped," Stuart said.
What Analysts and Amazon Are Saying
Bryan Silverthorn, director of AGI autonomy at Amazon, told attendees at VB Transform 2026 that the central obstacle is not intelligence but predictability. He argued that reliability must be measured across four dimensions: consistency, robustness, predictability, and safety. Amazon's approach involves sandboxed environments where agents propose changes for human review before implementation.
Gartner senior director analyst Shiva Varma warned that applying uniform governance across all AI agents raises failure rates. Gartner predicts that by 2027, 40% of companies will decommission agents because they failed to separate what an agent can do from the scope of access it was granted.
One Bright Spot: Evaluation Tooling Adoption
Even as trust in autonomous deployment fell, use of evaluation tooling climbed. OpenAI's native evaluation tools appeared in 59% of enterprise stacks in August, up from 31% in July. OpenAI was named the primary evaluation platform by 40% of respondents who answered that question. Organizations are not abandoning evaluation; they are investing more in it while pulling back on how much they trust its output.
The pattern that emerges from both surveys is straightforward: the competitive advantage in AI agents through 2027 will belong to companies that can get agents approved by risk, legal, and compliance teams, and keep them approved after go-live. Speed to deployment is no longer the differentiator. Verifiability is.
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
