Quick Facts
- Claude Sonnet 4.6 scored 79.6% on software coding benchmarks, nearly matching flagship Opus 4.6’s 80.8%
- Pricing remains at $3 per million input tokens versus Opus’s $15, delivering five times cost savings
- Enterprise accuracy jumped 15 percentage points across sectors, with legal tasks improving from 57% to 69%
Anthropic released Claude Sonnet 4.6 on Tuesday, delivering flagship model performance at one-fifth the cost of its premium Opus tier. The launch comes just 12 days after the company’s Opus 4.6 release.
The new model scored 79.6% on SWE-bench Verified, the industry standard for real-world software coding tasks. That nearly matches Opus 4.6’s 80.8% score while maintaining Sonnet’s $3 per million input token pricing versus Opus’s $15 rate.
“Performance that would have previously required reaching for an Opus-class model — including on real world, economically valuable office tasks — is now available with Sonnet 4.6,” Anthropic said in a blog post.
Box CEO Aaron Levie reported significant enterprise improvements across sectors. Public sector accuracy jumped from 77% to 88%, healthcare rose from 60% to 78%, and legal tasks improved from 57% to 69%.
“Box evaluated how Claude Sonnet 4.6 performs when tested on deep reasoning and complex agentic tasks across real enterprise documents,” said Ben Kus, Box’s chief technology officer. “It demonstrated significant improvements, outperforming Claude Sonnet 4.5 in heavy reasoning Q&A by 15 percentage points.”
The cost advantage proves transformational for enterprise deployments making millions of API calls daily. At scale, the difference between $15 and $3 per million tokens changes fundamental economics for AI-powered applications.
Anthropic’s computer use capabilities have nearly quintupled in 16 months. Sonnet 4.6 scored 72.5% on OSWorld-Verified benchmark, up from 14.9% when the feature launched in October 2024. The company graduated computer use from experimental status.
Claude Code, Anthropic’s developer tool launched publicly in May 2025, now generates over $2.5 billion in run-rate revenue. That figure more than doubled since January 2026. Business subscriptions quadrupled this year, with enterprise customers representing over half of Claude Code revenue.
The model includes a 1 million token context window in beta, double the previous Sonnet limit. Anthropic describes this as “enough to hold entire codebases, lengthy contracts, or dozens of research papers in a single request.”
User preference testing showed Sonnet 4.6 beat its predecessor 70% of the time. Users even preferred it over November’s flagship Opus 4.5 model 59% of the time.
Ramp data shows 1 in 5 businesses now pay for Anthropic, up from 1 in 25 a year ago. Among OpenAI customers, 79% also pay for Anthropic services.
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
