Skip to main navigation Skip to search Skip to main content

HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap

  • Wenxiang Lin
  • , Xinglin Pan
  • , Lin Zhang*
  • , Shaohuai Shi*
  • , Xuan Wang
  • , Xiaowen Chu
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • Hong Kong University of Science and Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The sparsely activated mixture-of-experts (MoE) transformer has become a common architecture for large language models (LLMs) due to its sparsity, which requires fewer computational demands while easily scaling the model size. In MoE models, each MoE layer requires to dynamically choose tokens to activate particular experts for computation while the activated experts may not be located in the same device or GPU as the token. However, this leads to substantial communication and load imbalances across all GPUs, which obstructs the scalability of distributed systems within a GPU cluster. To this end, we introduce HierMoE to accelerate the training of MoE models by two topology-aware techniques: 1) token deduplication to reduce the communication traffic, and 2) expert swap to balance the workloads among all GPUs. To enable the above two proposed approaches to be more general, we build theoretical models aimed at achieving the best token duplication and expert swap strategy under different model configurations and hardware environments. We implement our prototype HierMoE system atop Megatron-LM and conduct experiments on a 32-GPU cluster with DeepSeek-V3 and Qwen3-30B-A3B models. Experimental results show that our HierMoE achieves 1.55× to 3.32× faster communication and delivers 1.18× to 1.27× faster end-to-end training compared to state-of-the-art MoE training systems, Tutel-2DH, SmartMoE, and Megatron-LM.

Original languageEnglish
Title of host publicationINFOCOM 2026 - IEEE Conference on Computer Communications
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331549619
DOIs
StatePublished - 2026
Externally publishedYes
Event2026 IEEE Conference on Computer Communications, INFOCOM 2026 - Tokyo, Japan
Duration: 18 May 202621 May 2026

Publication series

NameProceedings - IEEE INFOCOM
ISSN (Print)0743-166X

Conference

Conference2026 IEEE Conference on Computer Communications, INFOCOM 2026
Country/TerritoryJapan
CityTokyo
Period18/05/2621/05/26

Keywords

  • Distributed Deep Learning
  • Expert Parallelism
  • Expert Swap
  • Mixture-of-Experts
  • Token Deduplication

Fingerprint

Dive into the research topics of 'HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap'. Together they form a unique fingerprint.

Cite this