Google AI released Multi-Token Prediction drafters for the Gemma 4 model family on May 6, 2026, using speculative decoding architecture to achieve up to 3x speedup in tokens-per-second for models like Gemma 4 31B without quality loss. The drafters address memory-bandwidth bottlenecks and are supported across LiteRT-LM, MLX, Hugging Face, and vLLM frameworks. The release follows Gemma 4's rapid adoption of over 60 million downloads in its first few weeks.
Aug 7, 2026 · 1 source
Aug 10, 2026 · 8 sources
Aug 7, 2026 · 5 sources
Aug 7, 2026 · 6 sources
Story comments
Loading comments…