Google has released Gemma 4 12B, an 11.95-billion-parameter open-source AI model designed to run locally on laptops with 16GB of RAM. Its novel "Unified" architecture processes audio and visual data without separate encoders, reducing latency. The model offers performance nearing Google's 26B model, supports a 256K token context window, and includes native tools for building autonomous agents for enterprise use.
Gemma 4 12B release
- ▪Gemma 4 12B is available for download on Hugging Face and Kaggle, and for use on Google AI Edge Gallery
- ▪Google released Gemma 4 12B, an 11.95-billion-parameter open-weights AI model
- ▪The Gemma 4 12B model is distributed under a permissive Apache 2.0 license
- ▪Gemma 4 12B is Google's first mid-sized model to feature native audio inputs
Encoder-free unified architecture
- ▪The encoder-free design reduces inference latency and memory requirements compared to traditional multimodal models
- ▪Gemma 4 12B features a "Unified" architecture that processes audio and visual data without separate, dedicated encoders
- ▪The model's vision encoder is replaced by a 35-million-parameter module using a single matrix multiplication
- ▪The audio encoder is eliminated entirely, with raw audio signals projected directly into the model's embedding space
Local execution capabilities
- ▪The model's media processing is limited to 30 seconds for audio inputs and 60 seconds for video at one frame per second
- ▪The downloadable model weights for Gemma 4 12B are just under 18GB in size
- ▪Gemma 4 12B is designed to run locally on laptops with 16GB of VRAM or unified memory
Performance benchmark results
- ▪On standard benchmarks, Gemma 4 12B's performance approaches that of Google's larger 26B Mixture-of-Experts model
- ▪The model supports a context window of up to 256,000 tokens, enabling it to process long documents or transcripts
- ▪Gemma 4 12B is the first model in its family to include Multi-Token Prediction (MTP) drafters out of the box for faster performance
Native agentic tool support
- ▪Gemma 4 12B has native support for function calling and system prompts, which are prerequisites for building autonomous agents
- ▪Google released an official Gemma Skills Repository to support developers building agents with Gemma models
- ▪The model includes a "thinking" mode that maps out step-by-step reasoning before it generates a response
Enterprise adoption use cases
- ▪Gemma 4 12B integrates with development frameworks including vLLM, SGLang, MLX, and llama.cpp
- ▪The model is positioned for cost-sensitive edge computing applications, such as in retail or for offline field service
- ▪Gemma 4 12B is intended for enterprises in regulated sectors like finance and healthcare that require strict data privacy
Story comments
Loading comments…