Quick Facts
- DeepSeek V4.1 Flash launched September 10, 2026, with off-peak cached-input pricing of $0.003 per million tokens, roughly 133 times cheaper than Claude Opus 5’s standard input rate.
- The 552B-parameter Mixture-of-Experts model beats GPT-5.6 Sol and Claude Opus 5 on four of five major agentic benchmarks, according to DeepSeek’s own model card.
- V4 Pro traffic will automatically route to V4.1 Flash after September 14, 2026, until a future V4.1 Pro release.
DeepSeek released V4.1 Flash on September 10, 2026, and the pricing alone commands attention. Off-peak cached-input tokens cost $0.003 per million. Cache misses run $0.15 per million. Output costs $0.60 per million. Peak rates, active during two narrow windows on weekday mornings UTC, are double those figures.
For U.S.-based engineering teams running standard business hours, nearly all API traffic falls into the off-peak window automatically. That makes the headline rate the practical rate for most American companies.
The gap versus Western competitors is large. OpenAI prices GPT-5.6 Sol at $4 per million standard input tokens and $20 per million output tokens. Anthropic charges $5 per million input and $25 per million output for Claude Opus 5. V4.1 Flash’s off-peak output rate of $0.60 per million is more than 40 times cheaper than Opus 5’s output pricing.
Compared with DeepSeek’s own prior Flash tier, the new model cuts cached-input costs by about 57%, uncached input by roughly 32%, and output by about 9%. Against the prior Pro tier, reductions reach 86% on cached input and 70% on output.
Architecture behind the cost structure
V4.1 Flash uses a 552B-parameter Mixture-of-Experts backbone but activates only 8 billion parameters per token during input processing and 16 billion per output token. DeepSeek calls this an asymmetric Causal Encoder-Decoder design. For agentic workloads where prompts are long and outputs are short, that asymmetry cuts compute on the expensive side of the transaction.
The model also carries an additional 196B-parameter Engram memory component and supports a 1 million-token context window with up to 384,000 output tokens. Its KV cache footprint runs roughly 890 bytes per token, about one-quarter of what V4 Flash required, and needs one-eighth the SSD storage of the prior generation.
DeepSeek’s API reports generation speed of 190.1 tokens per second, against a median of 64.9 tokens per second for open-weight models of comparable size. The model ships under the MIT license and handles image input natively.
Benchmark claims need context
According to DeepSeek’s model card, V4.1 Flash scores higher than GPT-5.6 Sol and Claude Opus 5 on four agentic benchmarks: DeepSWE v1.1 at 74.2 versus Sol’s 73.0, AutomationBench at 54.8 versus 45.8, Agent’s Last Exam at 31.8 versus 26.7, and CyberGym at 88.1 versus 84.5.
The model does not lead across every test. On Humanity’s Last Exam, V4.1 Flash scores 36.8% against Opus 5’s 56.3%. On GPQA Diamond, it scores 90.9%, behind Sol’s 94.1% and Opus 5’s 93.4%. V4 Pro also remains ahead on GPQA Diamond and on text-only reasoning subsets.
These numbers come from DeepSeek. Independent verification has not been published as of the release date.
What this means for enterprise buyers
The pricing structure rewards well-engineered agentic systems. Agents that reuse the same system prompt, tool list, and conversation history across turns benefit most from cache-hit pricing, where V4.1 Flash’s rate sits 50 times below the uncached price. Teams running high-volume pipelines with consistent context will see the largest cost reductions.
DeepSeek has scheduled V4 Pro traffic to redirect automatically to V4.1 Flash starting September 14, 2026. That means existing Pro customers get the new model without any code changes, and at substantially lower prices. A V4.1 Pro release is planned but has no confirmed date.
For founders and executives evaluating AI infrastructure costs, V4.1 Flash resets the competitive baseline on price for frontier-class performance on agentic tasks. Western providers have not announced matching price cuts.
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
