Skip to main navigation Skip to search Skip to main content

ProFetch: Accelerate Deep Recommendation System Training with Proactively Designed Data Layout and Dynamic Prefetching

  • Zhibing Liu
  • , Biyu Zhou*
  • , Weigang Zhang
  • , Xuehai Tang
  • , Ruixuan Li
  • , Songlin Hu
  • *Corresponding author for this work
  • University of Chinese Academy of Sciences
  • CAS - Institute of Information Engineering
  • Huazhong University of Science and Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Recommendation systems based on deep learning have played a vital role in today’s society. Since the embedding layers of recommendation models consume massive memory, current practice is to handle them on CPUs and use GPUs to speed up the training of the remaining parameters. In this hybrid training paradigm, data transfer between CPU and GPU becomes a bottleneck. While some cache-enhanced efforts have been proposed, the somewhat passive nature limits the throughput of training. In this paper, we propose ProFetch, a novel proactive cache prefetching method, to fully leverage the cache and accelerate training. Specifically, we observe varying degrees of overlap in the embedding parameters accessed between randomly shuffled mini-batches, and exploit this to propose a mini-batch layout strategy capable of radically reducing CPU-GPU data transfer during prefetching. We then propose an aggressive cache prefetching strategy that adaptively determines the content to prefetch at each step, maximizing the overlap in data transfer with GPU computing. We conduct prototype testing with open-sourced deep learning based recommendation models. Experimental results show that compared with representative methods, ProFetch significantly improves the training throughput of cache-enhanced hybrid recommendation systems, achieving speedups of up to 2.06X.

Original languageEnglish
Title of host publicationNeural Information Processing - 31st International Conference, ICONIP 2024, Proceedings
EditorsMufti Mahmud, Maryam Doborjeh, Kevin Wong, Andrew Chi Sing Leung, Zohreh Doborjeh, M. Tanveer
PublisherSpringer Science and Business Media Deutschland GmbH
Pages462-477
Number of pages16
ISBN (Print)9789819665877
DOIs
StatePublished - 2025
Externally publishedYes
Event31st International Conference on Neural Information Processing, ICONIP 2024 - Auckland, New Zealand
Duration: 2 Dec 20246 Dec 2024

Publication series

NameLecture Notes in Computer Science
Volume15290 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference31st International Conference on Neural Information Processing, ICONIP 2024
Country/TerritoryNew Zealand
CityAuckland
Period2/12/246/12/24

Keywords

  • cache prefetching
  • data layout
  • embedding layer
  • recommendation system
  • training

Fingerprint

Dive into the research topics of 'ProFetch: Accelerate Deep Recommendation System Training with Proactively Designed Data Layout and Dynamic Prefetching'. Together they form a unique fingerprint.

Cite this