Google AI released Multi-Token Prediction drafters for the Gemma 4 model family on May 6, 2026, using speculative decoding architecture to achieve up to 3x speedup in tokens-per-second for models like Gemma 4 31B without quality loss. The drafters address memory-bandwidth bottlenecks and are supported across LiteRT-LM, MLX, Hugging Face, and vLLM frameworks. The release follows Gemma 4's rapid adoption of over 60 million downloads in its first few weeks.
Sep 30, 2026 · 4 sources
Sep 28, 2026 · 7 sources
Sep 25, 2026 · 4 sources
Sep 30, 2026 · 3 sources
Story comments
Loading comments…