DeepSeek Made Its 75% Price Cut Permanent. SaaS Vendors Still Face a Margin Crisis.

Quick Facts

  • DeepSeek made its 75% promotional discount on V4-Pro API pricing permanent on May 22, 2026, setting the new list price at $0.435 per million input tokens and $0.87 per million output tokens.
  • The same 1 billion token monthly workload costs $522 on DeepSeek V4-Pro, versus $9,000 on Claude Opus 4.7 and $10,000 on GPT-5.5.
  • Enterprise SaaS vendors running AI agents are privately reporting negative gross margins on heavy users, as token amplification from agentic workflows outpaces per-token price cuts.

DeepSeek’s price cut was supposed to be temporary. On May 22, 2026, the Chinese AI lab announced it would not roll back the 75% discount on its V4-Pro API. The promotional rate became the permanent list price.

The new V4-Pro pricing sits at $0.435 per million input tokens on cache misses, $0.003625 per million on cache hits, and $0.87 per million output tokens. That makes it 7x cheaper on inputs and 17x cheaper on outputs than Anthropic’s Claude Sonnet or OpenAI’s GPT 5.5-Med.

For context: a workload of 800 million input tokens and 200 million output tokens per month costs $522 on V4-Pro. The same workload on Claude Opus 4.7 runs $9,000. On GPT-5.5, it hits $10,000.

Sanchit Vir Gogia, chief analyst and CEO of Greyhound Research, said the cut reflects engineering gains, not a sales play. “It is not a discount. It is an efficiency gain being passed through,” he said. V4-Pro runs at roughly 27% of the single-token compute that its predecessor required at 1 million token context, with KV cache memory dropping to about 10% of prior levels.

DeepSeek’s momentum on the open market is accelerating. On OpenRouter, DeepSeek V4-Flash ranked first globally with 3.43 trillion weekly token requests during the week of May 18-24. Total weekly requests across all DeepSeek models reached 5.74 trillion, surpassing Anthropic and Google for the second consecutive week. DeepSeek’s overall market share on OpenRouter climbed to 23.1%.

V4-Pro is now the world’s largest open-weight model at 1.6 trillion parameters. It scores 80.6% on the SWE-bench Verified coding-agent leaderboard and 87.5 on MMLU-Pro reasoning benchmarks.

Neil Shah, VP at Counterpoint Research, said V4-Pro has closed the performance gap on math and reasoning but still trails Western rivals on enterprise adoption, global support, IP provenance, and native hyperscaler integrations with AWS, Microsoft, and Google.

The Problem Cheaper Tokens Cannot Fix

Cheaper tokens help. They do not solve the core problem facing enterprise AI vendors: agentic workflows multiply token consumption far beyond what any pricing model anticipated.

A chatbot turns one user question into one model call. An agent turns it into a chain of planning, retrieval, tool use, verification, summarization, and follow-up decisions. A single user query can generate 700 times more tokens than a standard chat interaction.

Frontier inference costs are falling roughly 3x per year. But amplification is outrunning the cuts. A power user running 50 agent invocations per day on a $40 per seat plan can cost more in inference than the plan charges. Gross margins go negative.

Several enterprise SaaS vendors are now privately reporting exactly that outcome. The pattern mirrors findings from Bessemer’s Supernova cohort, where AI-agent adoption has shifted from a theoretical margin risk to a real P&L headwind. As one analysis framed it: the customers generating the most value are often the customers generating the highest inference costs.

The real-world numbers are stark. Uber exhausted its annual token budget in four months. Salesforce faces approximately $300 million in Anthropic costs this year alone.

Amit Jaju, senior managing director at Ankura Consulting, pointed to self-hosted deployments as one path out. “If a CIO can host DeepSeek V4-Pro on their own infrastructure, inference costs drop dramatically, and many projects that were previously uneconomical at scale become viable,” he said. That includes always-on copilots, bulk document review, code generation, and multi-agent workflows.

The pressure lands hardest on OpenAI and Anthropic, which are both absorbing the competitive pricing pressure from DeepSeek while their own enterprise customers run up inference bills that traditional SaaS subscription models were never designed to cover. OpenAI’s reported plan to give every Y Combinator startup $2 million in API credits reflects, as VentureBeat noted, what it now costs to run an AI-native company through its first year of product.

DeepSeek is also seeking its first external funding at a reported $44 billion valuation. The company runs V4-Pro on Huawei’s Ascend 950 chips and has declined to say whether improved chip supply contributed to making the price cut permanent.

Read more: DeepSeek cut prices 75%. The 100x problem remains

Get updates

Get curated daily technology news in your inbox.

Discover more from The SaaS Sentinel

Subscribe now to keep reading and get access to the full archive.

Continue reading