Ai2 Olmo-core 3 Scales Open MoE Training Beyond 2 Trillion Parameters
Ai2 opens Olmo-core 3, scaling mixture-of-experts training past 1 trillion parameters with 2.7 times throughput and 21% faster MXFP8 training performance.
Summary
On October 1, 2026, Ai2 released Olmo-core 3, an open training stack designed to take mixture-of-experts models into the trillion-parameter range and underpin the next Olmo. Expanding a benchmark model from 8 to 128 experts while selecting four per token held active parameters near 3.2 billion, increased total capacity from 4.6 billion to 47 billion, and reduced throughput by less than 5%. On eight NVIDIA B300 GPUs, that 47 billion parameter MoE reached 52,000 tokens per second per GPU, 2.7 times the earlier FSDP implementation's 19,400, after a DDP redesign kept experts resident on GPUs and routed data to them. OlmoE used 64 routed experts, while Olmo 3 was dense; NVIDIA Megatron-Core remains an established large-MoE alternative.
Expert parallelism, pipeline parallelism and a distributed optimizer divide experts, layers and optimizer state; rowwise expert parallelism, GPU-resident routing and grouped GEMM reduce overhead. On four B300s, MXFP8 delivered 21% more end-to-end throughput than BF16 and lowered peak active memory from 103 GiB to 95 GiB. A random-routing systems test of a 1.2 trillion parameter model, with 58.36 billion active per token across 512 B300s, peaked at 858 TFLOP/s/GPU but did not assess model quality; DeepEP v2 reached 2.38 trillion parameters only in a short capacity test. Ai2 also found token gerrymandering could disguise routing imbalance, lower expert learning rates did not help its tested family, matching shapes without matching values distorted GPU comparisons, and overlapping communication with computation sometimes slowed training. The GitHub code, technical report and interactive demo are open, while the planned MoE-based Olmo targets Ai2's largest dataset, longest context and highest capability yet.
Positives
- 128 experts provided 47 billion total parameters with about 3.2 billion active per token and less than 5% throughput loss.
- 52,000 tokens per second per GPU delivered roughly 2.7 times the earlier FSDP stack's throughput on eight NVIDIA B300 GPUs.
- MXFP8 increased throughput by 21% over BF16 while reducing peak active memory from 103 GiB to 95 GiB.
- 1.2 trillion parameters ran across 512 B300 GPUs at a peak 858 TFLOP/s/GPU.
- Open GitHub code, a technical report and an interactive demo let researchers adapt routing, parallelism and hardware support.
Risks & concerns
- 2.38 trillion parameters was reached only in a short DeepEP v2 capacity test, not a sustained training run.
- Random routing measured infrastructure performance at 1.2 trillion parameters but provided no evidence of trained model quality.
- Token gerrymandering allowed a routing score to improve while the real expert workload became less balanced.
- Lower expert learning rates failed to improve results in Ai2's tested model family.
- Overlapping communication and computation sometimes slowed execution, showing that individual optimizations can reduce end-to-end throughput.