The Allen Institute for AI has released OlmoCore 3, an open-source Mixture-of-Experts training framework that delivers a 2.7-fold throughput gain over prior implementations. By switching to distributed data parallelism and keeping experts resident on GPUs, the framework scales to trillion-parameter configurations. Alongside performance gains, the release openly documents training failure modes like 'token gerrymandering' to help academic and smaller labs audit their own pipelines.
Release and availability of OlmoCore 3
- ▪The Allen Institute for AI released OlmoCore 3, an open-source Mixture-of-Experts language model training framework, on October 1, 2026
- ▪OlmoCore 3 is released under the Apache 2.0 license and is available on GitHub and PyPI as ai2-olmo-core
Parallelism and data routing architecture
- ▪OlmoCore 3 utilizes expert parallelism, pipeline parallelism, and a distributed optimizer to distribute Mixture-of-Experts models across GPU clusters and reduce per-GPU memory demands
- ▪OlmoCore 3 transitioned from fully sharded data parallelism to distributed data parallelism, keeping experts resident on GPUs and routing training data to them to eliminate resharding overhead
Throughput and efficiency benchmarks
- ▪Enabling the MXFP8 reduced-precision format in OlmoCore 3 produced a 21% throughput gain over BF16 and reduced peak active memory from 103 GiB to 95 GiB on four B300 GPUs
- ▪In a benchmark on eight NVIDIA B300 GPUs, OlmoCore 3 achieved 52,000 tokens per second per GPU for a 47-billion-parameter MoE model, a 2.7-fold throughput increase over the prior implementation OlmoCore 2's 19,400 tokens
- ▪Scaling the expert pool from 8 to 128 experts in OlmoCore 3, while keeping active parameters at 3.2 billion, resulted in a training throughput drop of less than five percent
Large-scale parameter benchmarks
- ▪Testing OlmoCore 3 with the DeepEP v2 communication layer reached a configuration of approximately 2.38 trillion total parameters under artificial random routing
- ▪OlmoCore 3 reached a peak throughput of 858 TFLOPS per GPU during a 1.2-trillion-parameter configuration benchmark using 512 NVIDIA B300 GPUs under artificial random routing
Technical findings and training challenges
- ▪The OlmoCore 3 technical report noted that lowering learning rates for individual experts did not improve results, and overlapping communication and computation sometimes slowed end-to-end training
- ▪The Allen Institute for AI documented 'token gerrymandering,' a failure mode where the auxiliary loss metric for routing balance improves while the actual distribution of work across experts degrades
Allen Institute for AI initiatives
- ▪The Allen Institute for AI was selected in August 2025 by the National Science Foundation and NVIDIA for a $152 million initiative to build open scientific AI models
- ▪The Allen Institute for AI plans to train its next-generation OlmO model using a Mixture-of-Experts architecture on its largest dataset to date
Debatable claims
- ▪Publicly documenting unresolved training failures is necessary for scientific progress in AI
- ▪Open-source frameworks are capable of competing with proprietary industrial AI stacks
- ▪AI research organizations should open-source their underlying training infrastructure
Story comments
Loading comments…