New-ZZZ
RU / EN
LLM 1 October 2026

Olmo-core 3 Launches: New Infrastructure for Trillion-Parameter MoE Models

N
New-ZZZ desk
Hugging Face Blog · 14 hours ago

The release of Olmo-core 3 marks a major advancement in the infrastructure required for training massive, state-of-the-art Large Language Models (LLMs), specifically those utilizing the Mixture-of-Experts (MoE) architecture. This new framework is designed to overcome the significant computational and logistical hurdles associated with scaling MoE models into the trillion-parameter range, ensuring that advanced model development remains accessible and computationally efficient. The core challenge addressed by Olmo-core 3 is the inherent inefficiency that arises when scaling MoE models. While MoE models are inherently efficient because they only activate a small subset of parameters (the 'experts') for any given input token, the sheer size of the total model—which must still be stored and managed across a cluster of GPUs—introduces substantial communication and coordination overhead. As the number of potential experts grows, these communication costs can quickly erode the computational advantage that MoE models provide, making scaling difficult.

Olmo-core 3 tackles this scaling bottleneck by implementing a highly optimized and redesigned training system. A key demonstration of its capability is the ability to dramatically increase the expert pool size—in one benchmark, increasing it from 8 to 128 experts—while maintaining a stable, low number of active parameters per token, which remained fixed at approximately 3.2 billion parameters. Crucially, this massive increase in total model capacity, growing from 4.6 billion to 47 billion parameters, was achieved with a minimal degradation in training throughput, falling by less than 5%. This stability across exponential growth is the most critical breakthrough, enabling the industry to pursue models of unprecedented scale.

Technically, the framework represents a significant evolution from previous iterations. Earlier MoE implementations often relied on Fully Sharded Data Parallelism (FSDP), a method that gathers and reshard model weights for every small batch of training data. Olmo-core 3 makes a pivotal shift to a system based on Distributed Data Parallelism (DDP). This change is fundamental because DDP allows the specialized experts to remain resident directly on the GPUs. By keeping the experts localized, the system avoids the costly and time-consuming process of repeatedly gathering and redistributing the entire model's weights for every training step, thereby drastically improving efficiency and reducing latency.

Furthermore, Olmo-core 3 integrates and optimizes several advanced parallelization techniques to manage the model's vast size across multiple GPU clusters. These techniques include Expert Parallelism, which distributes the entire pool of experts across different GPUs so that no single GPU needs to store the full model. Pipeline Parallelism splits the model's sequential layers—the successive computational stages that transform the input data—across groups of GPUs, thereby reducing the memory burden on any single device. Finally, a Distributed Optimizer spreads the optimizer state (the auxiliary data used to calculate weight updates during training) across the GPUs, preventing any single GPU from needing to hold a full copy of this state. The synergy of these three parallelization methods allows the MoE architecture to scale effectively without requiring every piece of hardware to hold the entire model and its associated training state in memory.

Beyond the core parallelization strategies, Olmo-core 3 introduces several specialized optimizations that fine-tune the data flow and computation process. One such optimization is Rowwise Expert Parallelism, which is designed to place the routed data directly into the expert's input buffers. This minimizes the need for complex data rearrangement, which is a major source of computational waste. Another improvement is GPU-resident routing, which keeps the metadata related to data routing directly on the GPUs. This prevents the CPU from having to wait for this critical information to be copied back and forth, streamlining the entire workflow. Additionally, the framework incorporates Grouped GEMM (General Matrix Multiplication), a technique that groups multiple small expert computations together, allowing the GPUs to execute them in a highly efficient, batched manner, maximizing hardware utilization.

Another major enhancement is the support for MXFP8, a lower-precision number format. This format represents numerical values using fewer bits than standard formats like BF16. While using lower precision introduces potential conversion costs, the savings in computation and, critically, the reduction in the amount of data that must be moved between GPUs often outweigh these costs. In controlled benchmarks, enabling MXFP8 resulted in a substantial 21% increase in end-to-end training throughput compared to the BF16 baseline. Simultaneously, the peak active memory requirement dropped significantly, from 103 GiB to 95 GiB. This efficiency gain is primarily attributed to optimizing the feed-forward computation and the data movement between experts, rather than solely relying on the attention mechanism.

In practical terms, the performance gains are staggering. In a preliminary test conducted on eight NVIDIA B300 GPUs, the new Olmo-core 3 stack was able to process 52,000 tokens per second per GPU when running a 47-billion-parameter MoE. This represents a massive improvement—approximately 2.7 times faster—compared to the 19,400 tokens per second per GPU achieved using the framework's previous FSDP-based implementation. This dramatic increase in throughput solidifies Olmo-core 3 as a foundational tool that significantly lowers the barrier to entry for developing and training the next generation of massive, highly efficient AI models. The framework's continuous evolution, building upon earlier work like OlmoE (which used a 64-expert MoE) and Olmo 3 (which used a dense architecture), positions it as a comprehensive, scalable, and highly optimized training platform for the future of LLMs.

Why it matters

  • —It significantly lowers the computational barrier for developing trillion-parameter LLMs by optimizing MoE scaling.
  • —The shift from FSDP to DDP drastically improves training throughput and efficiency by keeping experts resident on GPUs.
  • —The integration of advanced techniques like MXFP8 and specialized parallelization methods maximizes hardware utilization and reduces memory footprint.

Key facts

  • Olmo-core 3 scales MoE training up to the trillion-parameter range.
  • It increased the expert pool from 8 to 128 while maintaining a stable active parameter count (~3.2B).
  • The new stack achieved 2.7x higher throughput (52,000 tokens/sec/GPU) compared to the previous implementation on a 47B MoE.
  • Support for MXFP8 boosted training throughput by 21% and reduced peak active memory from 103 GiB to 95 GiB.
Read the original →

The full text is in the original source. Here we provide a brief summary and key facts.

/ related