Skip to main navigation Skip to search Skip to main content

Demo: Split-and-Pipeline: Collaborative Large Model Inference on Edge Devices

  • Zuguang Li
  • , Dongyuan Ou
  • , Wen Wu*
  • , Songge Zhang
  • , Shaohua Wu
  • , Xuemin Shen
  • *Corresponding author for this work
  • Harbin Institute of Technology Shenzhen
  • Guangdong University of Technology
  • Peng Cheng Laboratory
  • University of Waterloo

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Deploying and executing large model inference on edge devices is challenging due to their limited computational power and memory resources. To address this challenge, we present a novel Split-and-Pipeline, a collaborative inference scheme that partitions a large model into multiple submodels and executes them across distributed edge devices in a pipelined manner. The scheme parallelizes data transfer across multiple CPU cores to avoid transmission bottlenecks. We build a real-world testbed using NVIDIA Jetson series edge devices to demonstrate the proposed scheme, achieving 1.2×–3.0× throughput improvement over state-of-the-art baselines.

Original languageEnglish
Title of host publicationACM MobiCom 2025 - Proceedings of the 2025 the 31st Annual International Conference on Mobile Computing and Networking
PublisherAssociation for Computing Machinery, Inc
Pages1207-1209
Number of pages3
ISBN (Electronic)9798400711299
DOIs
StatePublished - 21 Nov 2025
Externally publishedYes
Event31st Annual International Conference on Mobile Computing and Networking, ACM MobiCom 2025 - Hong Kong, China
Duration: 4 Nov 20258 Nov 2025

Publication series

NameACM MobiCom 2025 - Proceedings of the 2025 the 31st Annual International Conference on Mobile Computing and Networking

Conference

Conference31st Annual International Conference on Mobile Computing and Networking, ACM MobiCom 2025
Country/TerritoryChina
CityHong Kong
Period4/11/258/11/25

Keywords

  • Edge Networks
  • Large Models
  • Split-and-Pipeline Inference

Fingerprint

Dive into the research topics of 'Demo: Split-and-Pipeline: Collaborative Large Model Inference on Edge Devices'. Together they form a unique fingerprint.

Cite this