German AI startup Aleph Alpha has released Kolibri, an open-weight 78.1-billion-parameter Mixture-of-Experts language model optimized for German and English. Activating just 3.46 billion parameters per token, Kolibri targets regulated sectors like public administration and aviation. Developed in compliance with the EU AI Act and GDPR, the model supports a one-million-token context window and was trained on 24 trillion tokens using 768 NVIDIA B200 GPUs in Germany and Finland.
Development of Kolibri
- ▪Aleph Alpha's development teams in Germany controlled the entire pipeline for the Kolibri model, including data, architecture, training infrastructure, post-training, and evaluation
- ▪German AI startup Aleph Alpha released Kolibri, an open-weight Mixture-of-Experts language model built for German and English, under the Apache 2.0 license on Hugging Face
Deployment of Kolibri
- ▪Aleph Alpha designed the Kolibri language model for sovereign deployment in regulated sectors, specifically targeting public administration, industry, and aerospace or aviation
- ▪Aleph Alpha's Kolibri model can be served through vLLM using the aleph-alpha-inference package, which includes dedicated reasoning and tool-call parsers
Kolibri's model architecture
- ▪The Kolibri language model has 78.1 billion total parameters but activates only 3.46 billion parameters per token through a mixture-of-experts architecture
- ▪Kolibri's architecture features 50 transformer blocks with a model width of 2,560, utilizing 384 routed experts and 1 shared expert
- ▪Kolibri uses a hybrid attention mechanism with 40 sliding-window attention blocks and 10 full-attention blocks to keep the model's one-million-token context window computationally affordable
Context window configuration
- ▪Although Aleph Alpha's Kolibri supports a context window of over one million tokens, its default context length is 262,144 tokens unless overridden by the user
- ▪Aleph Alpha's Kolibri language model supports a context window of up to 1,048,576 tokens and allows users to set reasoning effort per request
Training data and vocabulary
- ▪Aleph Alpha utilized Chinese models to generate synthetic training data for the Kolibri language model
- ▪Kolibri's 128,000-token vocabulary was trained using UniBPE, achieving 4.90 bytes per token on German text and 4.58 bytes per token on English text
- ▪Kolibri was pre-trained on 24 trillion tokens, which included a dedicated German data pipeline making up 21.3% of the 24-trillion-token training total
Hardware and training infrastructure
- ▪Developed by Aleph Alpha, Kolibri's pre-training was executed on infrastructure consisting of 768 NVIDIA B200 GPUs located in Germany and Finland
- ▪The FP8 checkpoint of Kolibri is approximately 78 gigabytes and can run on a single NVIDIA B200, B300, or H200 GPU, or on two H100 SXM5 GPUs
Benchmark performance and evaluation
- ▪Kolibri decodes text approximately 2.7 times faster per GPU and scores 21.4 points higher in English than the predecessor model, Kolibri Origin
- ▪Kolibri achieved an overall score of 75.5 in English and 70.8 in German, outperforming comparable Mixture-of-Experts models evaluated by Aleph Alpha
- ▪In English benchmarks, Aleph Alpha's Kolibri scored 84.3 on GPQA Diamond, 96.9 on AIME 2025, and 96.0 on AIME 2026
Regulatory compliance
- ▪Aleph Alpha is a signatory of the EU General-Purpose AI Code of Practice, under which the Kolibri model was designed
- ▪To comply with the EU AI Act, GDPR, and the EU General-Purpose AI Code of Practice, Aleph Alpha designed Kolibri's data pipeline to redact personal data before training
Debatable claims
- ▪Strict compliance with European regulations limits the competitiveness of local AI models
- ▪Using Chinese models for synthetic training data compromises European AI sovereignty
Story comments
Loading comments…