Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
DeepSeek releases V4.1-Flash AI model with drastically reduced costs and memory usage
00

DeepSeek releases V4.1-Flash AI model with drastically reduced costs and memory usage

Sep 10, 2026

DeepSeek has launched its V4.1-Flash AI model, introducing a 552-billion-parameter Mixture-of-Experts architecture that dramatically reduces key-value cache memory requirements to one-quarter of its predecessor's footprint. By activating only 8 billion parameters during prefill and 16 billion during generation, the model achieves high-speed inference at a fraction of the cost of Western rivals. The release has triggered a mandatory API transition routing legacy V4-Pro traffic to V4.1-Flash by September 14, 2026, while sparking investor concerns over near-term semiconductor memory demand.

V4.1-Flash architecture improvements

  • ▪DeepSeek-V4.1-Flash was trained from scratch on a multimodal corpus containing 45 trillion tokens, with native image and text processing capabilities.
  • ▪DeepSeek-V4.1-Flash features a Causal Encoder-Decoder architecture that activates 8 billion parameters during input prefill and 16 billion parameters during text generation.
  • ▪DeepSeek launched DeepSeek-V4.1-Flash on September 10, 2026, as the smallest model in its new architecture family.
  • ▪DeepSeek-V4.1-Flash incorporates 196 billion N-gram parameters in a conditional memory module called Engram to store implicit knowledge and reduce compute requirements.
  • ▪DeepSeek-V4.1-Flash is built on a 552-billion-parameter Mixture-of-Experts framework, which is close to double the 284 billion parameters of the previous V4-Flash.

KV cache memory reduction

  • ▪DeepSeek-V4.1-Flash introduces Compressed Sparse Attention 2, which assigns Transformer layers to Full, Reindex, or Reuse modes to share cache and avoid redundant storage.
  • ▪DeepSeek-V4.1-Flash eliminates persistent sliding-window attention storage on SSDs, reducing persistent key-value cache storage requirements to roughly one-eighth of the previous generation.
  • ▪DeepSeek-V4.1-Flash reduces the global key-value cache footprint to 890 bytes per token, which is approximately one-quarter of the memory required by the previous V4-Flash.
  • ▪DeepSeek-V4.1-Flash stores its main global key-value cache in 4-bit floating-point format, which is a reduction from the 8-bit format used in the previous generation.

Agent workload cost economics

  • ▪DeepSeek-V4.1-Flash allows users to adjust its reasoning effort using an integer setting from 1 to 100, which trades off compute costs against accuracy.
  • ▪A VentureBeat Pulse Research survey from July 2026 found that only 47% of 170 surveyed enterprises rigorously track AI compute cost and ROI.
  • ▪DeepSeek-V4.1-Flash's architectural efficiency reduces the need for high-bandwidth memory, which contributed to a 3% drop in Samsung Electronics and SK Hynix shares on September 11, 2026.

API pricing structure

  • ▪DeepSeek-V4.1-Flash peak API rates are double its off-peak rates, rising to $0.006 per million cached input tokens, $0.15 per million uncached input tokens, and $0.60 per million output tokens.
  • ▪DeepSeek-V4.1-Flash off-peak API rates are priced at $0.003 per million tokens for cache hits, $0.075 per million tokens for cache misses, and $0.30 per million tokens for outputs.
  • ▪DeepSeek's peak API pricing windows run Monday through Friday from 01:00 to 04:00 UTC and from 06:00 to 10:00 UTC.

Mandatory endpoint transition timeline

  • ▪Starting September 14, 2026, at 04:00 UTC, DeepSeek will automatically reroute all API traffic directed to the deepseek-v4-pro endpoint to DeepSeek-V4.1-Flash.
  • ▪DeepSeek has retired the V4-Flash and V4-Flash-Vision-Exp models, temporarily routing their legacy API identifiers to DeepSeek-V4.1-Flash.

Benchmark performance comparisons

  • ▪DeepSeek-V4.1-Flash scored 90.6% on Terminal-Bench 2.1, outperforming OpenAI's GPT-5.6 Sol at 88.8% and Moonshot AI's Kimi K3 at 88.3%.
  • ▪DeepSeek's technical report notes that DeepSeek-V4.1-Flash's DeepSWE v1.1 score varied from 65.5% to 74.2% depending solely on the agent evaluation harness used.
  • ▪On the OpenDesign Arena benchmark, DeepSeek-V4.1-Flash scored 81.2 out of 100 on design tasks at a cost of $0.023 per design, compared to OpenAI's GPT-6 Astra which scored 82.7 at a cost of $1.61.
  • ▪On the DeepSWE v1.1 software engineering benchmark, DeepSeek-V4.1-Flash scored 74.2%, slightly ahead of Claude Opus 5 at 74.0% and GPT-5.6 Sol at 73.0%.

Debatable claims

  • ▪Using model distillation to train competitive AI models is ethically unacceptable
  • ▪Western enterprises should avoid integrating AI models from Chinese companies

12 sources

The-decoder
New Deepseek model V4.1-Flash cuts memory needs for AI agents
View source article
Reuters
China's DeepSeek launches V4.1-Flash model | Reuters
View source article
Thenews
DeepSeek launches V4.1-Flash model highlighting unmatched 400 plus token speeds
View source article
Decrypt
DeepSeek's New Model Nearly Matches GPT-6 Astra on Design—at 1.4% of the Cost - Decrypt
View source article
Cryptobriefing
DeepSeek launches V4.1-Flash model with 552B parameters and a million-token context window
View source article

Featured stories

View more in AI startups

DeepSeek releases software tools for Huawei AI chips to challenge Nvidia

Sep 30, 2026 · 4 sources

Tencent leases 100,000 Nvidia chips from Oracle in largest overseas cloud deal

Oct 1, 2026 · 4 sources

Google announces Gemini 4 Argon in limited release to cybersecurity partners

Sep 30, 2026 · 12 sources

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources

Story comments

Loading comments…

Related entities

China

Related Projects

DeepSeek

Topics

AI startupsAI tools & productsCompute, chips & AI infrastructureAI inference (scaling)AI foundation modelsChina tech & industrial strategy

Featured stories

View more in AI startups

DeepSeek releases software tools for Huawei AI chips to challenge Nvidia

Sep 30, 2026 · 4 sources

Tencent leases 100,000 Nvidia chips from Oracle in largest overseas cloud deal

Oct 1, 2026 · 4 sources

Google announces Gemini 4 Argon in limited release to cybersecurity partners

Sep 30, 2026 · 12 sources

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources