Quick Facts

  • 51% of organizations using AI experienced negative consequences, with most tied to AI inaccuracy, according to McKinsey’s 2025 survey
  • Andrej Karpathy’s ‘March of Nines’ concept shows moving from 90% to 99% AI reliability requires the same effort as reaching the initial 90%
  • Only 39% of organizations report enterprise-level EBIT impact from AI, with most seeing less than 5% contribution

Former OpenAI co-founder and Tesla Autopilot chief Andrej Karpathy has sounded a warning about AI reliability that new McKinsey data validates. His ‘March of Nines’ framework shows why 90% AI accuracy falls dangerously short of enterprise needs.

Karpathy’s concept, drawn from five years leading Tesla’s self-driving efforts, states that ‘every single nine is the same amount of work.’ Moving from 90% to 99% reliability requires equal effort as the initial demo that hit 90%.

The math is stark. An AI agent with 99% success rate per step has only 36.6% probability of completing a 100-step workflow successfully. Reliable agents need 99.99% reliability per step.

McKinsey Data Confirms Reliability Crisis

McKinsey’s 2025 global AI survey supports Karpathy’s warnings. Despite 88% of organizations now using AI in at least one business function, 51% experienced negative consequences. Nearly one-third of these problems stemmed from AI inaccuracy.

The survey found only 39% of organizations report any enterprise-level EBIT impact from AI. Among those seeing benefits, most attribute less than 5% of their organization’s EBIT to AI initiatives.

Two-thirds of organizations remain stuck in ‘pilot purgatory,’ unable to scale AI across the enterprise. McKinsey notes that high-performing AI adopters report more negative consequences because they deploy the technology in mission-critical contexts requiring sensitive monitoring.

Engineering Solutions Beyond Hype

Karpathy criticized current industry approaches in an October 2025 interview, saying ‘the models are not there. The industry is making too big of a jump and is trying to pretend like this is amazing, and it’s not—it’s slop.’

He estimates AGI remains about 10 years away, contrasting with OpenAI CEO Sam Altman’s prediction that AI will surpass human intelligence in any specialty by 2030.

The reliability gap requires disciplined engineering: bounded workflows, strict interfaces, resilient dependencies, and fast operational learning loops. Companies must define completion SLOs for top workflows and add contracts plus validators around every model output.

Market Reality Check

Karpathy’s assessment challenges the $14 trillion added to tech valuations since ChatGPT launched. The market has priced in a world where AI automates all digital work within two years. His ‘decade of agents’ timeline suggests this transformation will take much longer.

The former Tesla executive describes this period as when the cost-performance curve of AI tokens intersects with human labor costs across sectors. This starts with low-stakes tasks before gradually moving to higher-reliability applications.

Read more: Karpathy’s March of Nines shows why 90% AI reliability isn’t even close to enough

This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.