Skip to main navigation Skip to search Skip to main content

HyDLR: Load-Aware Dynamic Rescheduling for Deep Learning Hybrid Deployment

  • Harbin Institute of Technology
  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Resource contention, driven by traffic surges from online services, presents a significant challenge in hybrid clusters where latency-sensitive and best-effort deep learning tasks are colocated. To address this, we propose HyDLR, a dynamic, loadaware hybrid deployment scheduling method that dynamically reallocates offline tasks to ensure Quality of Service (QoS) for online services while enhancing overall resource utilization. The bursty nature and stringent QoS demands of online tasks, coupled with the fluctuating resource footprints of offline tasks, can lead to severe resource pressure on nodes and undermine system stability. HyDLR first designs a load-aware rescheduling policy that dynamically identifies resource hotspots by monitoring metrics such as CPU satisfaction degree, memory, and GPU memory utilization. It then leverages eviction and task migration to optimize workload distribution. Furthermore, a two-stage filtering algorithm, guided by a multi-objective optimization model, targets system-wide load balancing and minimal rescheduling overhead. By incorporating a dynamically adjusted priority queue and a cost-feedback mechanism, HyDLR improves scheduling efficiency without compromising stability. Experimental results demonstrate that HyDLR significantly reduces the frequency of task migrations while achieving a well-balanced system load. The rate of cascading rescheduling events is kept below 3%, demonstrating superior performance over existing approaches. This work offers an effective solution for resource management in complex, hybrid deployment scenarios, laying a foundation for more efficient data center scheduling and demonstrating strong potential for practical adoption.

Original languageEnglish
Title of host publicationProceedings of 2025 IEEE 31st International Conference on Parallel and Distributed Systems, ICPADS 2025
PublisherIEEE Computer Society
ISBN (Electronic)9798331549015
DOIs
StatePublished - 2025
Externally publishedYes
Event31st IEEE International Conference on Parallel and Distributed Systems, ICPADS 2025 - Hefei, China
Duration: 14 Dec 202517 Dec 2025

Publication series

NameProceedings of the International Conference on Parallel and Distributed Systems - ICPADS
ISSN (Print)1521-9097

Conference

Conference31st IEEE International Conference on Parallel and Distributed Systems, ICPADS 2025
Country/TerritoryChina
CityHefei
Period14/12/2517/12/25

Keywords

  • Cloud Data Center
  • DNN Task
  • Resource Allocation
  • Task Scheduling

Fingerprint

Dive into the research topics of 'HyDLR: Load-Aware Dynamic Rescheduling for Deep Learning Hybrid Deployment'. Together they form a unique fingerprint.

Cite this