Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
DeepSeek releases V4.1-Flash AI model with drastically reduced costs and memory usage
00

DeepSeek releases V4.1-Flash AI model with drastically reduced costs and memory usage

Sep 10, 2026

DeepSeek has launched its V4.1-Flash AI model, introducing a 552-billion-parameter Mixture-of-Experts architecture that dramatically cuts key-value cache memory requirements to one-quarter of its predecessor's footprint. By activating only 8 billion parameters during prefill and 16 billion during generation, the model achieves high-speed inference at a fraction of the cost of US rivals. The aggressive pricing and efficiency gains have pressured local competitors and triggered stock declines for South Korean memory chipmakers Samsung and SK Hynix.

V4.1-Flash model release

  • ▪DeepSeek-V4.1-Flash is built on a 552-billion-parameter Mixture-of-Experts framework, activating 8 billion parameters during input prefill and 16 billion parameters during output generation.
  • ▪DeepSeek released the weights for DeepSeek-V4.1-Flash on Hugging Face under the permissive MIT license, requiring approximately 475 GiB of storage across 48 safetensors shards.
  • ▪DeepSeek-V4.1-Flash features native multimodal visual understanding and supports a context window of up to 1 million tokens.
  • ▪DeepSeek launched the DeepSeek-V4.1-Flash artificial intelligence model on September 10, 2026, as the smallest model in its new architecture family.
  • ▪DeepSeek-V4.1-Flash incorporates a 196-billion-parameter conditional memory module called Engram, which uses N-gram lookup tables to store implicit knowledge and reduce compute requirements.

KV cache memory reduction

  • ▪DeepSeek-V4.1-Flash stores its main global KV cache in 4-bit floating-point (FP4) precision, which nearly halves the memory footprint compared to the 8-bit floating-point (FP8) format used in V4.
  • ▪DeepSeek-V4.1-Flash achieves KV cache reduction by splitting its 40 Transformer layers into an asymmetric 20-layer causal encoder and a 20-layer decoder.
  • ▪DeepSeek-V4.1-Flash reduces persistent SSD storage requirements to one-eighth of the previous generation by offloading sliding-window attention state to a temporary DRAM pool and using Bounded Replay to reconstruct state.
  • ▪DeepSeek-V4.1-Flash reduces the active key-value (KV) cache memory footprint to approximately 890 bytes per token, which is one-quarter of the memory required by the previous V4-Flash model.
  • ▪DeepSeek-V4.1-Flash introduces an adjustable reasoning effort setting from 1 to 100, where higher settings improve benchmark accuracy but consume up to 2.5 times more output tokens.

API pricing structure

  • ▪DeepSeek priced the V4.1-Flash API off-peak rates at $0.003 per million tokens for cached inputs, $0.15 per million for uncached inputs, and $0.60 per million for outputs, with rates doubling during weekday peak hours.
  • ▪DeepSeek's off-peak output rate of $0.30 per million tokens for V4.1-Flash is more than 83 times cheaper than Anthropic's Claude Opus-5 output rate of $25 per million tokens.
  • ▪DeepSeek retired its V4-Flash and V4-Flash-Vision-Exp models on September 10, 2026, and announced that all API traffic to the deepseek-v4-pro endpoint will automatically route to V4.1-Flash starting September 14, 2026.

Performance benchmark results

  • ▪DeepSeek's technical report notes that the same V4.1-Flash checkpoint scored between 65.5% and 74.2% on DeepSWE v1.1 depending solely on which agent evaluation harness wrapped the model.
  • ▪DeepSeek-V4.1-Flash scored 74.2% on the DeepSWE v1.1 software engineering benchmark at maximum reasoning effort, narrowly outperforming Anthropic's Claude Opus-5 at 74.0% and OpenAI's GPT-5.6 Sol at 73.0%.
  • ▪DeepSeek-V4.1-Flash scored 90.6 on Terminal-Bench 2.1, surpassing OpenAI's GPT-5.6 Sol at 88.8, Moonshot AI's Kimi K3 at 88.3, and DeepSeek's own V4 Pro at 87.9.
  • ▪On the OpenDesign Arena benchmark, DeepSeek-V4.1-Flash scored 81.2 out of 100 on real-world design tasks, reaching 98% of OpenAI's GPT-6 Astra score of 82.7 while costing 1.4% of Astra's price.

Market impact on competitors

  • ▪Following the launch of DeepSeek-V4.1-Flash, shares of South Korean memory chipmakers Samsung Electronics and SK Hynix both fell more than 3 percent on September 11, 2026, due to investor concerns over reduced hardware demand.
  • ▪The launch of DeepSeek-V4.1-Flash triggered an 8 percent drop in the Hong Kong shares of Chinese AI competitors MiniMax Group and Z.AI, while Alibaba shares slid more than 2 percent.
  • ▪The launch of DeepSeek-V4.1-Flash coincided with reports that DeepSeek has hired CITIC Securities to prepare for an initial public offering on Shanghai's tech-focused STAR Market.

Debatable claims

  • ▪Western enterprises should avoid integrating AI models from Chinese companies
  • ▪AI agent skills look great in benchmarks but fall apart under realistic conditions according to researchers
  • ▪AI software efficiency gains will ultimately reduce global demand for semiconductor hardware

12 sources

Cryptobriefing
DeepSeek launches V4.1-Flash model with 552B parameters and a million-token context window
View source article
Scmp
DeepSeek says new Flash AI model beats Kimi K3 on cyber, coding benchmarks
View source article
The-decoder
New Deepseek model V4.1-Flash cuts memory needs for AI agents
View source article
Analyticsindiamag
AIM — India's Leading AI & Data Science Media Platform
View source article
Thenews
DeepSeek launches V4.1-Flash model highlighting unmatched 400 plus token speeds
View source article

Featured stories

View more in AI research & benchmarks

Nvidia CEO Jensen Huang says AGI has arrived, congratulates OpenAI on GPT-6 Astra

Sep 6, 2026 · 4 sources

Nvidia launches PAIR tool for home AI computing

Sep 3, 2026 · 3 sources

DeepSeek launches hiring spree for 150 engineers to overhaul strained infrastructure

Sep 8, 2026 · 1 source

IFM releases K2 Horizon family of six open models from 0.9B to 375B parameters

Sep 7, 2026 · 2 sources

Story comments

Loading comments…

Related entities

China

Related Projects

DeepSeek

Topics

AI research & benchmarksCompute, chips & AI infrastructureLarge language models (LLMs)AI tools & productsAI foundation modelsAI inference (scaling)

Featured stories

View more in AI research & benchmarks

Nvidia CEO Jensen Huang says AGI has arrived, congratulates OpenAI on GPT-6 Astra

Sep 6, 2026 · 4 sources

Nvidia launches PAIR tool for home AI computing

Sep 3, 2026 · 3 sources

DeepSeek launches hiring spree for 150 engineers to overhaul strained infrastructure

Sep 8, 2026 · 1 source

IFM releases K2 Horizon family of six open models from 0.9B to 375B parameters

Sep 7, 2026 · 2 sources