Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
DeepSeek releases V4-Flash-0731 with major performance gains
00

DeepSeek releases V4-Flash-0731 with major performance gains

Jul 31, 2026

DeepSeek released V4-Flash-0731 on July 31, 2026, a major upgrade to its AI model with significant gains in agentic tasks. The 284B-parameter model now scores 50 on the Artificial Analysis index, nearly matching OpenAI's GPT-5.6 Luna at a 60% lower cost. The update, achieved via post-training, is available through a public beta API and as MIT-licensed weights on Hugging Face.

V4-Flash-0731 release details

  • ▪The model's weights are available on Hugging Face under an MIT license, allowing for on-premise commercial deployment
  • ▪DeepSeek released the V4-Flash-0731 model on Hugging Face and moved its V4-Flash API into public beta on July 31, 2026
  • ▪The 0731 release is an update to the preview version from April 2026, with gains from re-post-training rather than architectural changes

API pricing structure

  • ▪Per-task costs for V4-Flash-0731 are about 60% lower than for OpenAI's GPT-5.6 Luna
  • ▪The API offers a 98% cache discount, with cache hits costing $0.0028 per 1M input tokens
  • ▪The V4-Flash API is priced at $0.14 per 1M input tokens and $0.28 per 1M output tokens

Self-hosting requirements

  • ▪A full-precision version of the model can be served on a single 4xGB300 node using DeepSeek's vLLM example
  • ▪Self-hosting the model requires significant memory, with a 3-bit quantized build needing around 110 GB of combined RAM and VRAM

Model architecture specifications

  • ▪V4-Flash is a 284-billion-parameter Mixture-of-Experts (MoE) model that activates 13 billion parameters per token
  • ▪The model features a 1-million-token context window
  • ▪The checkpoint includes the DSpark speculative decoding module, bringing the total parameter count in the Hugging Face repo to 304B

Benchmark performance results

  • ▪DeepSeek's reported benchmarks, including an 82.7 on Terminal Bench 2.1, were run using an unreleased "minimal mode" of its testing harness
  • ▪The model's score on the GDPval benchmark, which tests complex office work, increased from 1,189 to 1,559 Elo points
  • ▪On the Artificial Analysis Intelligence Index, V4-Flash-0731 scored 50 points, just one point behind OpenAI's GPT-5.6 Luna

Serving configuration options

  • ▪The DSpark speculative decoding module can be enabled in vLLM and is reported to speed up per-user generation by 60-85%
  • ▪For agentic tasks, DeepSeek recommends setting temperature to 1.0 and top_p to 0.95

3 sources

Bloomberg
DeepSeek Unveils Public Beta API for Flagship AI Model
View source article
Marktechpost
DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major Agentic and Coding Gains
View source article
The-decoder
New Deepseek Flash model matches OpenAI's GPT-5.6 Luna at roughly 60 percent lower cost
View source article

Featured stories

View more in AI startups

OpenAI revenue hits $70 billion annualized rate as ChatGPT reaches 1.2 billion weekly users

Sep 29, 2026 · 13 sources

FTC opens investigation into OpenAI and Anthropic over consumer protection

Sep 30, 2026 · 7 sources

OpenAI launches Dots, always-on AI agents that work across 4,000+ apps

Sep 29, 2026 · 14 sources

OpenAI announces Codex cloud environments, Decisions API and Ultrafast tier at DevDay 2026

Sep 29, 2026 · 7 sources

Story comments

Loading comments…

Related Projects

DeepSeek

Topics

AI startupsLarge language models (LLMs)OpenAIAI coding assistantsAI agentsAI research & benchmarks

Featured stories

View more in AI startups

OpenAI revenue hits $70 billion annualized rate as ChatGPT reaches 1.2 billion weekly users

Sep 29, 2026 · 13 sources

FTC opens investigation into OpenAI and Anthropic over consumer protection

Sep 30, 2026 · 7 sources

OpenAI launches Dots, always-on AI agents that work across 4,000+ apps

Sep 29, 2026 · 14 sources

OpenAI announces Codex cloud environments, Decisions API and Ultrafast tier at DevDay 2026

Sep 29, 2026 · 7 sources