Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
Google Releases Multi-Token Prediction Drafters for Gemma 4 with 3x Speedup
00

Google Releases Multi-Token Prediction Drafters for Gemma 4 with 3x Speedup

May 7, 2026

Google AI released Multi-Token Prediction drafters for the Gemma 4 model family on May 6, 2026, using speculative decoding architecture to achieve up to 3x speedup in tokens-per-second for models like Gemma 4 31B without quality loss. The drafters address memory-bandwidth bottlenecks and are supported across LiteRT-LM, MLX, Hugging Face, and vLLM frameworks. The release follows Gemma 4's rapid adoption of over 60 million downloads in its first few weeks.

Memory-bandwidth bottleneck

  • ▪During standard Large Language Model inference, the processor spends the majority of its operational time moving billions of parameters from Video RAM to the compute units just to generate a single token
  • ▪Multi-Token Prediction drafters for Gemma 4 specifically target memory-bandwidth limitations that often hinder performance on consumer-grade hardware
  • ▪Standard Large Language Model inference processes are often limited by the speed at which data can be moved rather than by the processor's calculation speed
  • ▪The memory-bandwidth bottleneck leads to under-utilized compute resources and high latency, particularly on consumer-grade hardware where memory speeds may not match the demands of high-parameter models

Multi-token prediction architecture

  • ▪The Multi-Token Prediction drafter is a specialized, smaller model that can perform predictions in significantly less time than it takes the larger target model to process a single token
  • ▪Google released Multi-Token Prediction drafters for the Gemma 4 family of open models on May 6, 2026
  • ▪The Multi-Token Prediction drafter utilizes idle compute cycles to predict several future tokens simultaneously
  • ▪Multi-Token Prediction drafters for Gemma 4 allow models like Gemma 4 31B to achieve up to a 3x speedup in tokens-per-second

Speculative decoding mechanics

  • ▪The Multi-Token Prediction drafters for Gemma 4 achieve speed increases with no degradation in reasoning logic or output quality
  • ▪Google's Multi-Token Prediction drafters use a specialized speculative decoding architecture to address latency bottlenecks in AI inference
  • ▪Speculative decoding with Multi-Token Prediction drafters decouples token generation from verification by pairing a heavy target model with a lightweight drafter model
  • ▪In speculative decoding with Multi-Token Prediction drafters, once the drafter has proposed a sequence of tokens, the target model performs a verification step and accepts multiple tokens at once if the predictions are accurate

Framework compatibility

  • ▪The broad framework compatibility of Multi-Token Prediction drafters allows developers to implement faster inference on diverse hardware, from Apple Silicon via MLX to cloud-based deployments via vLLM
  • ▪Multi-Token Prediction drafters for Gemma 4 are supported across LiteRT-LM, MLX, and Hugging Face frameworks

Hardware deployment options

  • ▪Google designed Multi-Token Prediction drafters for Gemma 4 to improve responsiveness across mobile devices, developer workstations, and the cloud
  • ▪Multi-Token Prediction drafters for Gemma 4 make high-capability open models more responsive and viable for real-time applications without requiring specialized, high-end enterprise hardware for every use case
  • ▪Multi-Token Prediction drafters for Gemma 4 enhance responsiveness for developers working on mobile devices, workstations, and cloud environments

Gemma 4 adoption metrics

  • ▪The Gemma 4 family of models has seen over 60 million downloads in its first few weeks
  • ▪Google describes Gemma 4 as delivering unprecedented intelligence-per-parameter
  • ▪The release of Multi-Token Prediction drafters for Gemma 4 follows the model's rapid adoption of over 60 million downloads

2 sources

Aitoolly
Google Gemma 4 MTP Drafters: 3x Faster AI Inference Speed | AIToolly
View source article
Marktechpost
Google AI Releases Multi-Token Prediction (MTP) Drafters for Gemma 4: Delivering Up to 3x Faster Inference Without Quality Loss
View source article

Featured stories

View more in Large language models (LLMs)

DeepSeek releases software tools for Huawei AI chips to challenge Nvidia

Sep 30, 2026 · 4 sources

AMD acquires AI startup World Labs founded by Fei-Fei Li for $8.2 billion

Sep 28, 2026 · 7 sources

AI models Astra and Claude Opus crack unsolved World War II Enigma messages

Sep 25, 2026 · 4 sources

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources

Story comments

Loading comments…

Related Projects

Google AIGoogle

Topics

Large language models (LLMs)AI research & benchmarksSpeculative decodingAI inference (scaling)Open-source AICompute, chips & AI infrastructure

Featured stories

View more in Large language models (LLMs)

DeepSeek releases software tools for Huawei AI chips to challenge Nvidia

Sep 30, 2026 · 4 sources

AMD acquires AI startup World Labs founded by Fei-Fei Li for $8.2 billion

Sep 28, 2026 · 7 sources

AI models Astra and Claude Opus crack unsolved World War II Enigma messages

Sep 25, 2026 · 4 sources

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources