DeepSeek launches DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal model that adds visual processing to its low-cost V4 Flash architecture. While DeepSeek claims performance close to Anthropic's active Opus 4.8 model, benchmark data shows the new model wins on only three of eleven tests, trailing significantly on repository-scale tasks. Commercially, the model remains highly competitive due to DeepSeek's extremely low API pricing compared to US rivals.
V4 Flash Vision Exp release
- ▪DeepSeek released an experimental multimodal artificial intelligence model named DeepSeek-V4-Flash-Vision-Exp on August 21, 2026, making it available to developers via the DeepSeek API platform.
- ▪DeepSeek-V4-Flash-Vision-Exp adds visual processing capabilities, enabling the system to process and execute tasks based on still photographs and screen captures, to DeepSeek's text-only V4 Flash model.
- ▪DeepSeek released version 0.1.1 of its agent harness on August 21, 2026, providing out-of-the-box support for the DeepSeek-V4-Flash-Vision-Exp model.
Opus 4.8 benchmark comparison
- ▪DeepSeek-V4-Flash-Vision-Exp outperformed Anthropic's Opus 4.8 model on three out of eleven benchmarks published by DeepSeek, specifically winning on DeepSWE, Agents' Last Exam, and ZeroBench.
- ▪Anthropic's Opus 4.8 model, which was released in May 2026, remains active and fully supported with no planned retirement earlier than May 2027.
- ▪DeepSeek-V4-Flash-Vision-Exp trailed Anthropic's Opus 4.8 model on eight of eleven benchmarks, including a 12-point deficit on NL2Repo and an 8.1-point deficit on DSBench-Hard.
Multimodal agent performance gains
- ▪DeepSeek disclosed in a benchmark footnote that the text-only V4 Flash model's lower scores on multimodal tests occurred because the text-only model ignores visual elements.
- ▪On multimodal agent benchmarks, DeepSeek-V4-Flash-Vision-Exp scored 36.5 on ApexBench and 27.3 on Agents' Last Exam, representing a performance increase over the text-only V4 Flash model.
V4 Flash technical architecture
- ▪The DeepSeek V4 Flash model utilizes Hybrid Cache Attention and Compressed Sparse Attention to compress its KV cache, which reduces prompt processing compute requirements by 73% according to DeepMind.
- ▪The foundational DeepSeek V4 Flash model is a mixture-of-experts model containing 284 billion parameters, structured across multiple neural networks of 13 billion parameters each.
- ▪DeepSeek trained the foundational V4 Flash model on 32 trillion tokens using an algorithm called Muon to accelerate the calibration of the model's hidden layers.
Text capability improvements
- ▪DeepSeek-V4-Flash-Vision-Exp outperformed the text-only V4 Flash model on six out of seven text benchmarks, including Toolathlon-Verified, DeepSWE, and DSBench-Hard.
- ▪DeepSeek-V4-Flash-Vision-Exp scored 75.3 on the Cybergym security benchmark, representing a 1.4-point decrease compared to the text-only V4 Flash model's score of 76.7.
Story comments
Loading comments…