Quick Facts
- 85% of enterprises that experienced an AI failure are deploying without human review or building toward it, per July 2026 VB Pulse research.
- 49% of surveyed companies said an AI agent that cleared internal testing later caused a customer-visible problem, nearly unchanged from 50% the prior month.
- Only 5% of respondents fully trust the automated evaluations driving those deployment decisions.
Enterprises that have already been burned by AI failures in production are not slowing down. They are speeding up. That is the central finding from VentureBeat’s July 2026 Agentic Reliability and Evaluations Tracker, which surveyed 108 organizations with 100 or more employees.
Among companies that had experienced a testing miss, 85% were already deploying AI agents without human review or actively building toward that model. The same group was also investing most aggressively in people-centered review workflows, with 38% naming it their fastest-growing budget area. They are moving humans away from the deployment gate while funding them downstream.
The Numbers Behind the Gap
Across all respondents, 66% already permit some production deployment without human review or plan to within 12 months. Only 5% say they fully trust the automated evaluations that govern those decisions. That gap between autonomy and assurance is widening.
Trust in automated evaluation is rising. In July, 13% of respondents said they trust it, up from 5% the month before. But the failure rate is not falling. Nearly half of respondents, 49%, said an AI agent that passed company testing later caused a problem visible to customers. In June, that figure was 50%.
Twenty-four percent said this outcome had occurred more than once. The tools are getting more trusted. The results are not getting better.
What Companies Are Actually Watching
Among 106 valid responses, monitoring priorities reveal a structural blind spot. Twenty-six percent used inline quality assertions to check whether live output was correct. Another 26% tracked transaction traces such as token usage and raw inputs and outputs. Twenty-four percent focused on gateway metrics like latency and cost.
Grouped together, roughly half of respondents monitored whether an agent was functioning. Just over a quarter automatically monitored whether its output was correct. Companies are measuring the pipes, not the water.
The Confidence Inversion
The survey exposes a striking pattern around experience and trust. Among enterprises that had never suffered a false-confidence failure, 24% fully trust automated evaluation. Among those that had experienced a failure, only 4% do. The research notes that trust in this market is a function of exposure, not evidence.
That inversion helps explain the budget trends. People-centered review workflows received the most frequent mention for increased investment at 31%, narrowly ahead of production observability at 30%. Automated evaluation pipelines ranked third at 19%. Only 6% of respondents said their reliability and evaluation budget was not growing at all.
The Klarna Warning
The report arrives against a backdrop of high-profile reversals. Klarna became the most referenced cautionary example after announcing in 2024 that AI had effectively replaced approximately 700 customer service agents. The company then rebuilt human capacity through 2025 and into 2026.
CEO Sebastian Siemiatkowski acknowledged the shift, saying customers wanted access to a human and that offering one was critical from a brand perspective. Analyst Gary Marcus labeled the arc from AI triumphalism to reversal the Klarna Effect. A separate survey from Orgvue and Forrester found 55% of executives now openly regret replacing human workers with AI. Robert Half reported that 32% of U.S. hiring managers eliminated a role due to AI and later rehired for the same position.
The Vendor Market
Specialist evaluation platforms are gaining ground. The share of enterprises running no dedicated evaluation tooling dropped from 17% to 12% month over month. DeepEval reached 17% share, Braintrust nearly doubled to 15%, and OpenAI’s native evals led at 18%. Selection criteria shifted sharply: ease of integration overtook cost as the top factor, jumping from 27% to 39% of respondents.
The strategic picture is straightforward. Enterprises are buying evaluation tools faster, trusting them more, deploying agents with less human oversight, and still failing at nearly the same rate. The companies that have already paid the cost of a public failure are not pulling back. They are betting that better tooling will close a gap that better tooling has not yet closed.
Read more: 85% of companies burned by an AI mistake are racing to cut the humans who might catch the next one
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
