Skip to main navigation Skip to search Skip to main content

Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules

  • Xinglin Pan*
  • , Wenxiang Lin
  • , Shaohuai Shi
  • , Xiaowen Chu
  • , Weinong Sun
  • , Bo Li
  • *Corresponding author for this work
  • The Hong Kong University of Science and Technology (Guangzhou)
  • Hong Kong Baptist University
  • School of Computer Science and Technology, Harbin Institute of Technology
  • Hong Kong University of Science and Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Sparsely-activated Mixture-of-Expert (MoE) layers have found practical applications in enlarging the model size of large-scale foundation models, with only a sub-linear increase in computation demands. Despite the wide adoption of hybrid parallel paradigms like model parallelism, expert parallelism, and expert-sharding parallelism (i.e., MP+EP+ESP) to support MoE model training on GPU clusters, the training efficiency is hindered by communication costs introduced by these parallel paradigms. To address this limitation, we propose Parm, a system that accelerates MP+EP+ESP training by designing two dedicated schedules for placing communication tasks. The proposed schedules eliminate redundant computations and communications and enable overlaps between intra-node and inter-node communications, ultimately reducing the overall training time. As the two schedules are not mutually exclusive, we provide comprehensive theoretical analyses and derive an automatic and accurate solution to determine which schedule should be applied in different scenarios. Experimental results on an 8-GPU server and a 32-GPU cluster demonstrate that Parm outperforms the state-of-the-art MoE training system, DeepSpeed-MoE, achieving 1.13× to 5.77× speedup on 1296 manually configured MoE layers and approximately 3× improvement on two real-world MoE models based on BERT and GPT-2.

Original languageEnglish
Title of host publicationIEEE INFOCOM 2024 - IEEE Conference on Computer Communications
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages1880-1889
Number of pages10
ISBN (Electronic)9798350383508
DOIs
StatePublished - 2024
Externally publishedYes
Event43rd IEEE Conference on Computer Communications, INFOCOM 2024 - Vancouver, Canada
Duration: 20 May 202423 May 2024

Publication series

NameProceedings - IEEE INFOCOM
ISSN (Print)0743-166X

Conference

Conference43rd IEEE Conference on Computer Communications, INFOCOM 2024
Country/TerritoryCanada
CityVancouver
Period20/05/2423/05/24

Keywords

  • Distributed Training
  • Large Language Models
  • Mixture-of-Experts
  • Task Scheduling

Fingerprint

Dive into the research topics of 'Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules'. Together they form a unique fingerprint.

Cite this