Skip to main navigation Skip to search Skip to main content

Improving the Downstream Performance of Mixture-of-Experts Transformers via Weak Vanilla Transformers

  • Faculty of Computing, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Recently, mixture-of-experts (MoE) Transformers have garnered increased attention for their advantages in model capacity and computational efficiency. However, studies have indicated that MoE models excel in pre-training but fail to translate those gains to downstream tasks, which diminishes their practical value. To explain this issue, we propose that the pre-training performance and transfer capability of a model jointly determine its downstream task performance. More specifically, we argue that MoE models exhibit weaker transfer capability compared to vanilla models, resulting in inferior performance on downstream tasks. To address this issue, we introduce the concept of transfer capability distillation, positing that although vanilla models have weaker performance, they are effective teachers of transfer capability. The MoE models—guided by vanilla models—can achieve both strong pre-training performance and transfer capability, ultimately enhancing their performance on downstream tasks. We developed a specific distillation method and conducted experiments using the BERT architecture. Experimental results show significant improvements in the downstream performance of MoE models, and further evidence supports the concept of transfer-capability distillation. Finally, we attempt to interpret transfer capability distillation and provide insights from the perspective of model features.

Original languageEnglish
Article number4256
JournalElectronics (Switzerland)
Volume14
Issue number21
DOIs
StatePublished - Nov 2025
Externally publishedYes

Keywords

  • mixture-of-experts (MoE)
  • pre-trained language models
  • transfer capability

Fingerprint

Dive into the research topics of 'Improving the Downstream Performance of Mixture-of-Experts Transformers via Weak Vanilla Transformers'. Together they form a unique fingerprint.

Cite this