Skip to main navigation Skip to search Skip to main content

DeepFill: Accelerating MLLM Training by Filling Bubbles with Frozen Encoders

  • Harbin Institute of Technology Shenzhen

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The training of multimodal large language models (MLLMs) has emerged as a crucial area in artificial intelligence, aiming to integrate diverse modalities such as text and images into a unified framework. Due to the vast number of model parameters, MLLM training often employs techniques such as pipeline parallelism (PP) and the Zero Redundancy Optimizer (ZeRO) to address memory limitations. However, existing frameworks, such as DeepSpeed, fail to fully utilize GPU resources, resulting in significant idle time when distributing the training task across multiple devices. To mitigate this issue, we introduce DeepFill, a novel framework designed to enhance MLLM training efficiency. First, DeepFill separates the inference component of the frozen modality encoders from the training process of the main large language model (LLM), assigning them to distinct execution streams. Second, DeepFill leverages the idle GPU time in PP and ZeRO and precomputes the encoders within a single training step (PP bubbles) or between two consecutive steps (ZeRO bubbles). Our experiments on two open-source MLLMs demonstrate that DeepFill significantly improves the training throughput by up to 1.08× for PP, 1.11× for ZeRO and 1.18× for their hybrid optimization, closely aligning with theoretical expectations.

Original languageEnglish
Title of host publicationAlgorithms and Architectures for Parallel Processing - 25th International Conference, ICA3PP 2025, Proceedings
EditorsShadi Ibrahim, Thomas Rauber, Huazhong Liu
PublisherSpringer Science and Business Media Deutschland GmbH
Pages233-245
Number of pages13
ISBN (Print)9789819584109
DOIs
StatePublished - 2026
Externally publishedYes
Event25th International Conference on Algorithms and Architectures for Parallel Processing, ICA3PP 2025 - Zhengzhou, China
Duration: 30 Oct 20252 Nov 2025

Publication series

NameLecture Notes in Computer Science
Volume16386 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference25th International Conference on Algorithms and Architectures for Parallel Processing, ICA3PP 2025
Country/TerritoryChina
CityZhengzhou
Period30/10/252/11/25

Keywords

  • Distributed Training
  • Multimodel Large Language Model

Fingerprint

Dive into the research topics of 'DeepFill: Accelerating MLLM Training by Filling Bubbles with Frozen Encoders'. Together they form a unique fingerprint.

Cite this