Google has released DiffusionGemma, an experimental AI model that uses text diffusion to generate text up to 4x faster than traditional models. Designed for interactive local workflows, the 26B parameter model generates entire blocks of text at once, achieving over 1000 tokens/second on an NVIDIA H100. While its output quality is lower than standard Gemma 4, it excels at non-linear tasks like code infilling.
DiffusionGemma model release
- ▪The model is built upon Google's Gemma 4 family and Gemini Diffusion research
- ▪For applications that demand maximum quality, Google recommends deploying its standard autoregressive Gemma 4 models
- ▪DiffusionGemma is released under a permissive Apache 2.0 license, with its experimental model weights available on Hugging Face
- ▪On June 10, 2026, Google introduced DiffusionGemma, an experimental open model that explores text diffusion for faster text generation
Text diffusion architecture
- ▪The model features bi-directional attention, allowing every token in a generated block to attend to all others
- ▪Instead of generating text token-by-token, the model generates entire blocks of 256 tokens simultaneously with each forward pass
- ▪The text diffusion process starts with a canvas of random placeholder tokens and iteratively refines them into a final output
- ▪DiffusionGemma is a 26B Mixture of Experts (MoE) model that activates only 3.8B parameters during inference
GPU inference speed benchmarks
- ▪On a single NVIDIA H100 GPU, the model can generate over 1000 tokens per second
- ▪The model's speed advantage is strongest for local and low-concurrency inference at low-to-medium batch sizes on a single accelerator
- ▪DiffusionGemma can deliver up to 4x faster text generation on dedicated GPUs compared to typical autoregressive models
Hardware requirements for deployment
- ▪When quantized, DiffusionGemma can fit within the 18GB VRAM limits of high-end dedicated consumer GPUs
- ▪Unified-memory architectures like those in Apple Silicon Macs may not see the same acceleration as they are often memory-bandwidth-bound
- ▪Google and NVIDIA collaborated to optimize the model for hardware including GeForce RTX 5090 and 4090 GPUs, as well as Hopper and Blackwell systems
Interactive editing use cases
- ▪DiffusionGemma is designed for researchers and developers exploring speed-critical, interactive local workflows like in-line editing and rapid iteration
- ▪The model's architecture provides advantages for non-linear domains such as code infilling, amino acid sequences, or mathematical graphs
- ▪As an example, the company Unsloth fine-tuned DiffusionGemma to play Sudoku, a task that is difficult for autoregressive models
Developer integration tools
- ▪Fine-tuning can be explored with tools such as Google's Hackable Diffusion, Unsloth, and NVIDIA NeMo
- ▪The model can be served using tools like MLX, vLLM (with integration supported by Red Hat), and Hugging Face Transformers
- ▪Official support for llama.cpp is expected to arrive soon after the model's release
Story comments
Loading comments…