Skip to main navigation Skip to search Skip to main content

Cola: Cross-Processor Operator Parallelism for Asynchronous Deep Learning Inference

  • Changyao Lin
  • , Zhenming Chen
  • , Ziyang Zhang
  • , Jie Liu*
  • *Corresponding author for this work
  • School of Computer Science and Technology, Harbin Institute of Technology
  • Ltd
  • Polytechnic University of Milan

Research output: Contribution to journalArticlepeer-review

Abstract

Multi-task inference, as a prevalent inference paradigm nowadays, requires deploying multiple deep learning models on the hardware platform to concurrently process inference tasks. Modern platforms are typically equipped with various heterogeneous processors, such as CPU-GPU platform. To reduce resource contention and improve quality of service in the multi-task scenario, existing work has studied cross-processor inference at the fine-grained operator-level. However, it lacks specific optimizations for asynchronous multi-task inference systems. In such systems, tasks arrive dynamically, leading to diverse inference progress for each model. This renders offline optimization strategies based solely on the original computation graph suboptimal or even ineffective. Therefore, we propose a novel framework, Cola, to address the cross-processor operator scheduling for asynchronous tasks. Cola introduces intermediate representation to abstract and simplify such dynamic scheduling problem, considering the impact of task arrival patterns on the inference progress, and employs an efficient two-phase search algorithm. We implemented and validated Cola on a real-world case of intelligent steel structure manufacturing. Cola outperforms the state-of-the-art cross-processor operator scheduling framework in both throughput and resource utilization with highly acceptable runtime overhead.

Original languageEnglish
Pages (from-to)4346-4362
Number of pages17
JournalIEEE Transactions on Mobile Computing
Volume25
Issue number3
DOIs
StatePublished - 2026

Keywords

  • Edge computing
  • cross-processor parallelism
  • multi-task inference
  • operator scheduling

Fingerprint

Dive into the research topics of 'Cola: Cross-Processor Operator Parallelism for Asynchronous Deep Learning Inference'. Together they form a unique fingerprint.

Cite this