DeepSeek has launched its V4.1-Flash AI model, introducing a 552-billion-parameter Mixture-of-Experts architecture that dramatically cuts key-value cache memory requirements to one-quarter of its predecessor's footprint. By activating only 8 billion parameters during prefill and 16 billion during generation, the model achieves high-speed inference at a fraction of the cost of US rivals. The aggressive pricing and efficiency gains have pressured local competitors and triggered stock declines for South Korean memory chipmakers Samsung and SK Hynix.
V4.1-Flash model release
- ▪DeepSeek-V4.1-Flash is built on a 552-billion-parameter Mixture-of-Experts framework, activating 8 billion parameters during input prefill and 16 billion parameters during output generation.
- ▪DeepSeek released the weights for DeepSeek-V4.1-Flash on Hugging Face under the permissive MIT license, requiring approximately 475 GiB of storage across 48 safetensors shards.
- ▪DeepSeek-V4.1-Flash features native multimodal visual understanding and supports a context window of up to 1 million tokens.
- ▪DeepSeek launched the DeepSeek-V4.1-Flash artificial intelligence model on September 10, 2026, as the smallest model in its new architecture family.
- ▪DeepSeek-V4.1-Flash incorporates a 196-billion-parameter conditional memory module called Engram, which uses N-gram lookup tables to store implicit knowledge and reduce compute requirements.
KV cache memory reduction
- ▪DeepSeek-V4.1-Flash stores its main global KV cache in 4-bit floating-point (FP4) precision, which nearly halves the memory footprint compared to the 8-bit floating-point (FP8) format used in V4.
- ▪DeepSeek-V4.1-Flash achieves KV cache reduction by splitting its 40 Transformer layers into an asymmetric 20-layer causal encoder and a 20-layer decoder.
- ▪DeepSeek-V4.1-Flash reduces persistent SSD storage requirements to one-eighth of the previous generation by offloading sliding-window attention state to a temporary DRAM pool and using Bounded Replay to reconstruct state.
- ▪DeepSeek-V4.1-Flash reduces the active key-value (KV) cache memory footprint to approximately 890 bytes per token, which is one-quarter of the memory required by the previous V4-Flash model.
- ▪DeepSeek-V4.1-Flash introduces an adjustable reasoning effort setting from 1 to 100, where higher settings improve benchmark accuracy but consume up to 2.5 times more output tokens.
API pricing structure
- ▪DeepSeek priced the V4.1-Flash API off-peak rates at $0.003 per million tokens for cached inputs, $0.15 per million for uncached inputs, and $0.60 per million for outputs, with rates doubling during weekday peak hours.
- ▪DeepSeek's off-peak output rate of $0.30 per million tokens for V4.1-Flash is more than 83 times cheaper than Anthropic's Claude Opus-5 output rate of $25 per million tokens.
- ▪DeepSeek retired its V4-Flash and V4-Flash-Vision-Exp models on September 10, 2026, and announced that all API traffic to the deepseek-v4-pro endpoint will automatically route to V4.1-Flash starting September 14, 2026.
Performance benchmark results
- ▪DeepSeek's technical report notes that the same V4.1-Flash checkpoint scored between 65.5% and 74.2% on DeepSWE v1.1 depending solely on which agent evaluation harness wrapped the model.
- ▪DeepSeek-V4.1-Flash scored 74.2% on the DeepSWE v1.1 software engineering benchmark at maximum reasoning effort, narrowly outperforming Anthropic's Claude Opus-5 at 74.0% and OpenAI's GPT-5.6 Sol at 73.0%.
- ▪DeepSeek-V4.1-Flash scored 90.6 on Terminal-Bench 2.1, surpassing OpenAI's GPT-5.6 Sol at 88.8, Moonshot AI's Kimi K3 at 88.3, and DeepSeek's own V4 Pro at 87.9.
- ▪On the OpenDesign Arena benchmark, DeepSeek-V4.1-Flash scored 81.2 out of 100 on real-world design tasks, reaching 98% of OpenAI's GPT-6 Astra score of 82.7 while costing 1.4% of Astra's price.
Market impact on competitors
- ▪Following the launch of DeepSeek-V4.1-Flash, shares of South Korean memory chipmakers Samsung Electronics and SK Hynix both fell more than 3 percent on September 11, 2026, due to investor concerns over reduced hardware demand.
- ▪The launch of DeepSeek-V4.1-Flash triggered an 8 percent drop in the Hong Kong shares of Chinese AI competitors MiniMax Group and Z.AI, while Alibaba shares slid more than 2 percent.
- ▪The launch of DeepSeek-V4.1-Flash coincided with reports that DeepSeek has hired CITIC Securities to prepare for an initial public offering on Shanghai's tech-focused STAR Market.
Debatable claims
- ▪Western enterprises should avoid integrating AI models from Chinese companies
- ▪AI agent skills look great in benchmarks but fall apart under realistic conditions according to researchers
- ▪AI software efficiency gains will ultimately reduce global demand for semiconductor hardware
Story comments
Loading comments…