Abstract
Multi-task inference, as a prevalent inference paradigm nowadays, requires deploying multiple deep learning models on the hardware platform to concurrently process inference tasks. Modern platforms are typically equipped with various heterogeneous processors, such as CPU-GPU platform. To reduce resource contention and improve quality of service in the multi-task scenario, existing work has studied cross-processor inference at the fine-grained operator-level. However, it lacks specific optimizations for asynchronous multi-task inference systems. In such systems, tasks arrive dynamically, leading to diverse inference progress for each model. This renders offline optimization strategies based solely on the original computation graph suboptimal or even ineffective. Therefore, we propose a novel framework, Cola, to address the cross-processor operator scheduling for asynchronous tasks. Cola introduces intermediate representation to abstract and simplify such dynamic scheduling problem, considering the impact of task arrival patterns on the inference progress, and employs an efficient two-phase search algorithm. We implemented and validated Cola on a real-world case of intelligent steel structure manufacturing. Cola outperforms the state-of-the-art cross-processor operator scheduling framework in both throughput and resource utilization with highly acceptable runtime overhead.
| Original language | English |
|---|---|
| Pages (from-to) | 4346-4362 |
| Number of pages | 17 |
| Journal | IEEE Transactions on Mobile Computing |
| Volume | 25 |
| Issue number | 3 |
| DOIs | |
| State | Published - 2026 |
Keywords
- Edge computing
- cross-processor parallelism
- multi-task inference
- operator scheduling
Fingerprint
Dive into the research topics of 'Cola: Cross-Processor Operator Parallelism for Asynchronous Deep Learning Inference'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver