Quick Facts
- DeepSeek released DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal model that adds image understanding to its V4 Flash text model.
- The model costs $0.22 per million input tokens, compared to pricing that can run workloads costing $10,000 per month on Claude Opus 4.8 down to roughly $120 to $600 per month on DeepSeek’s V4 Pro.
- DeepSeek also released version 0.1.1 of its Harness framework and a new Files API alongside the model launch.
DeepSeek released a new experimental model on August 21 that processes both text and images, claiming performance close to Anthropic’s Claude Opus 4.8 at a steep price discount. The model, called DeepSeek-V4-Flash-Vision-Exp, is a vision-enabled variant of V4 Flash, the smaller and faster model DeepSeek released in April alongside V4 Pro.
The release extends DeepSeek’s V4 series into multimodal territory without launching a separate product. The model handles text and image inputs natively through a single architecture, processing JPEG, PNG, GIF, and WebP files.
Architecture and Performance
DeepSeek-V4-Flash-Vision-Exp is a sparse Mixture-of-Experts model with 13 billion active parameters out of 284 billion total. It carries a context window of 1,048,576 tokens and supports up to 384,000 completion tokens. Each image input is capped at 384 tokens, with larger images normalized by default to roughly 800 by 800 pixels.
DeepSeek published benchmark comparisons against Claude Opus 4.8 across multiple tests. The new model scored 27.3 on Agents’ Last Exam against Opus 4.8’s 25.7, and 35.0 on ZeroBench against Opus 4.8’s 34.0. On ApexBench, DeepSeek scored 36.5 at Pass@1 compared to Opus 4.8’s 39.4. On Terminal Bench 2.1, the model scored 83.9 against Opus 4.8’s 85.0.
Wider gaps remain on some tests. On NL2Repo, DeepSeek scored 57.7 against Opus 4.8’s 69.7. The company also acknowledged underperformance on Cybergym, which evaluates software vulnerability discovery. According to reporting from SiliconAngle, DeepSeek’s model wins three of eleven benchmarks in the published comparison table. DeepSeek ran evaluations using its internal Harness Minimal Mode, and scores have not been independently verified.
Pricing and API Access
The model is priced at $0.22 per million input tokens and $0.66 per million output tokens during off-peak hours. During peak hours, defined as 01:00 to 04:00 and 06:00 to 10:00 UTC, the rate rises to $0.44 per million input tokens and $1.32 per million output tokens. Image processing carries no price premium over the base Flash model.
Developers access the model through the same API endpoint used for V4 Flash. The model is compatible with OpenAI’s Chat Completions and Responses APIs and Anthropic’s Messages endpoint, which lowers the switching cost for teams already working with those interfaces.
New Developer Tools
DeepSeek released version 0.1.1 of its Harness framework alongside the model, with built-in support for V4-Flash-Vision-Exp. The company also launched a Files API that lets developers upload images once and reference them by file ID across multiple requests at no additional cost.
The Files API supports storage of up to 25 GiB and 10,000 files per user. File expiration can be set from one hour to 30 days at the time of upload. Files with no expiration setting are stored permanently. Developers can also send images via Base64 encoding or publicly accessible URLs up to 32 MiB. An optional detail field downscales images to 512 by 512 pixels when fine visual detail is not needed, reducing token consumption.
What It Means for Enterprise Buyers
The release adds competitive pressure to Anthropic at the high end of the model market. DeepSeek’s pricing gap remains significant. A workload costing $10,000 per month on Opus 4.8 could run for $120 to $600 per month on DeepSeek’s V4 Pro, according to the company’s own estimates. For software teams running high-volume document processing, screenshot analysis, or agent workflows, the cost differential is material.
At launch, the model is available only through DeepSeek’s paid API. Open-source availability has not been announced.
Read more: DeepSeek debuts multimodal language model competitive with Opus 4.8
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
