Quick Facts
- DeepSeek V4 Flash completed only 53.8% of complex real-world agent tasks across 240 runs in independent testing by Composio.
- DeepSeek raised API prices up to 1,100% for V4 Flash and V4 Pro, effective Aug. 16, with new peak and off-peak billing tiers.
- The official benchmark harness DeepSeek uses to report scores has not been released publicly, making third-party reproduction of its numbers impossible.
DeepSeek’s V4 Flash has earned strong praise from developers since its rollout, topping AI leaderboards and posting dramatic benchmark gains. But independent testing shows the model struggles when put to work on real tasks — and the company has sharply raised prices for access.
Composio tested V4 Flash across eight agent harnesses on 30 complex workflows, each involving multiple SaaS applications. The apps included Airtable, Gmail, Google Calendar, Google Sheets, GitHub, Slack, and PostHog. Each task had a 900-second limit, and results were independently verified. Across 240 total runs, harnesses completed 129 successfully — a 53.8% pass rate.
Only six of the 30 workflows were completed successfully by every harness. Pi led all harnesses at 66.7%, completing 20 of 30 tasks. Prime Agent reported 62.5%. OMP finished 17 tasks at 56.7%. Claude Code, Codex, and DeepAgents each hit 53.3%. Hermes finished 50%, and OpenCode came in last at 46.7%.
The spread across harnesses running the same model points to a core issue. Tool configuration, caching behavior, retry logic, and provider stack all affected outcomes substantially. Raw model capability alone does not determine whether agents succeed in enterprise workflows.
The Benchmark Gap
On paper, V4 Flash’s gains look impressive. The model scored 54.4 on DeepSWE and 82.7 on Terminal Bench 2.1, up from 7.3 and 61.8 on the same benchmarks for the previous version. Its GDPval-AA v2 Elo score rose from 1,189 to 1,559. DeepSeek achieved those gains entirely through post-training, not a new model architecture.
V4 Flash uses a 284-billion parameter Mixture-of-Experts design with 13 billion activated parameters and a 1-million-token context window. The V4 family launched April 24, 2026, under an MIT license. The July 31 build, V4-Flash-0731, carries the same weights and architecture with only its post-training updated.
The benchmark numbers, though, carry a significant caveat. DeepSeek reports its official scores using its own internal harness, which has not been released publicly. No third party can reproduce those numbers. A separate audit of V4 Pro by yage.ai in June 2026 scored it at 8% pass@1 on DeepSWE, against 70% for GPT-5.5 and 54% for Anthropic’s Opus 4.7 — a sharp contrast to the 54.4 DeepSeek’s harness reports for Flash 0731 on the same benchmark.
Prices Rise Sharply
DeepSeek’s price increases took effect at 16:00 UTC on Aug. 16. The company introduced a peak and off-peak billing structure, with peak hours defined as 01:00 to 04:00 UTC and 06:00 to 10:00 UTC. Off-peak rates are set at half the peak rates, but every off-peak rate still exceeds the previous flat rate.
V4 Flash output tokens now cost $1.32 per million at peak, up from a flat $0.28 per million — a 371% increase. Off-peak output comes in at $0.66 per million, still more than double the old rate. V4 Pro output tokens rise to $3.96 per million at peak and $1.98 per million off-peak, compared to the prior $0.87 per million flat rate.
Input token pricing also climbed. V4 Flash cache-miss input tokens rose to $0.44 per million at peak from $0.14. V4 Pro cache-miss input tokens increased to $1.32 per million at peak from $0.435. The steepest increases exceed 1,100% on certain token types during peak hours.
DeepSeek said the changes are meant to “allocate resources more reasonably” and push workloads toward less congested periods. For developers running high-volume production workloads during peak hours, the cost structure has changed materially. Off-peak scheduling will become a critical consideration for teams that built cost models around DeepSeek’s previous flat pricing.
The combination of a real-world performance gap and a significant price increase gives enterprise buyers two reasons to reassess assumptions about V4 Flash. Leaderboard position tells one story. What happens across 240 production runs tells another.
Read more: DeepSeek’s top-ranked V4 Flash stumbles on real agent tasks as its prices surge
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
