TY - GEN
T1 - HyDLR
T2 - 31st IEEE International Conference on Parallel and Distributed Systems, ICPADS 2025
AU - Wang, Desheng
AU - Sun, Xiao
AU - Si, Shuo
AU - Chen, Sichao
AU - Zhang, Weizhe
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Resource contention, driven by traffic surges from online services, presents a significant challenge in hybrid clusters where latency-sensitive and best-effort deep learning tasks are colocated. To address this, we propose HyDLR, a dynamic, loadaware hybrid deployment scheduling method that dynamically reallocates offline tasks to ensure Quality of Service (QoS) for online services while enhancing overall resource utilization. The bursty nature and stringent QoS demands of online tasks, coupled with the fluctuating resource footprints of offline tasks, can lead to severe resource pressure on nodes and undermine system stability. HyDLR first designs a load-aware rescheduling policy that dynamically identifies resource hotspots by monitoring metrics such as CPU satisfaction degree, memory, and GPU memory utilization. It then leverages eviction and task migration to optimize workload distribution. Furthermore, a two-stage filtering algorithm, guided by a multi-objective optimization model, targets system-wide load balancing and minimal rescheduling overhead. By incorporating a dynamically adjusted priority queue and a cost-feedback mechanism, HyDLR improves scheduling efficiency without compromising stability. Experimental results demonstrate that HyDLR significantly reduces the frequency of task migrations while achieving a well-balanced system load. The rate of cascading rescheduling events is kept below 3%, demonstrating superior performance over existing approaches. This work offers an effective solution for resource management in complex, hybrid deployment scenarios, laying a foundation for more efficient data center scheduling and demonstrating strong potential for practical adoption.
AB - Resource contention, driven by traffic surges from online services, presents a significant challenge in hybrid clusters where latency-sensitive and best-effort deep learning tasks are colocated. To address this, we propose HyDLR, a dynamic, loadaware hybrid deployment scheduling method that dynamically reallocates offline tasks to ensure Quality of Service (QoS) for online services while enhancing overall resource utilization. The bursty nature and stringent QoS demands of online tasks, coupled with the fluctuating resource footprints of offline tasks, can lead to severe resource pressure on nodes and undermine system stability. HyDLR first designs a load-aware rescheduling policy that dynamically identifies resource hotspots by monitoring metrics such as CPU satisfaction degree, memory, and GPU memory utilization. It then leverages eviction and task migration to optimize workload distribution. Furthermore, a two-stage filtering algorithm, guided by a multi-objective optimization model, targets system-wide load balancing and minimal rescheduling overhead. By incorporating a dynamically adjusted priority queue and a cost-feedback mechanism, HyDLR improves scheduling efficiency without compromising stability. Experimental results demonstrate that HyDLR significantly reduces the frequency of task migrations while achieving a well-balanced system load. The rate of cascading rescheduling events is kept below 3%, demonstrating superior performance over existing approaches. This work offers an effective solution for resource management in complex, hybrid deployment scenarios, laying a foundation for more efficient data center scheduling and demonstrating strong potential for practical adoption.
KW - Cloud Data Center
KW - DNN Task
KW - Resource Allocation
KW - Task Scheduling
UR - https://www.scopus.com/pages/publications/105032479438
U2 - 10.1109/ICPADS67057.2025.11323176
DO - 10.1109/ICPADS67057.2025.11323176
M3 - 会议稿件
AN - SCOPUS:105032479438
T3 - Proceedings of the International Conference on Parallel and Distributed Systems - ICPADS
BT - Proceedings of 2025 IEEE 31st International Conference on Parallel and Distributed Systems, ICPADS 2025
PB - IEEE Computer Society
Y2 - 14 December 2025 through 17 December 2025
ER -