Skip to main navigation Skip to search Skip to main content

Dataset Pruning: Reducing Training Data by Examining SGD-Influence

  • Shuo Yang
  • , Yucheng Huang
  • , Zeke Xie
  • , Ping Li
  • , Min Xu
  • , Liqiang Nie*
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • The University of Hong Kong
  • The Hong Kong University of Science and Technology (Guangzhou)
  • VecML Inc.
  • University of Technology Sydney

Research output: Contribution to journalArticlepeer-review

Abstract

The great success of deep learning relies heavily on ever-increasing amounts of training data, which incurs enormous computational and infrastructural costs. This raises crucial questions: Does all training data contribute equally to a model's performance? How much does each individual training sample or sub-training set affect the model's generalization, and how can we construct the smallest proxy training set without significantly sacrificing performance? To address these questions, we propose dataset pruning, an optimization-based sample selection method that (1) examines the influence of removing particular training samples on the model's generalization ability with high computational efficiency and theoretical guarantees, and (2) constructs the smallest subset of training data that yields a strictly constrained generalization gap. The empirically observed generalization gap achieved by dataset pruning is largely consistent with our theoretical expectations. To estimate each sample's influence, we develop an SGD-Influence method that tracks parameter changes during stochastic gradient descent, effectively overcoming the convexity and optimality assumptions required by traditional influence function approaches. Complementing this, our distributed discrete optimization partitions the dataset into manageable buckets, allowing for efficient sample selection without sacrificing quality. Extensive experiments on three datasets of varying scale and complexity demonstrate that our approach not only aligns well with theoretical expectations but also outperforms state-of-the-art methods. Notably, compared to our previous work, the proposed method achieves a 61.26% reduction in computational cost while delivering higher accuracy.

Original languageEnglish
JournalIEEE Transactions on Pattern Analysis and Machine Intelligence
DOIs
StateAccepted/In press - 2026
Externally publishedYes

Keywords

  • data pruning
  • efficient deep learning
  • influence function

Fingerprint

Dive into the research topics of 'Dataset Pruning: Reducing Training Data by Examining SGD-Influence'. Together they form a unique fingerprint.

Cite this