Skip to main navigation Skip to search Skip to main content

ClusMatch: Improving Deep Clustering by Unified Positive and Negative Pseudo-Label Learning

  • Jianlong Wu
  • , Zihan Li
  • , Wei Sun
  • , Jianhua Yin*
  • , Liqiang Nie*
  • , Zhouchen Lin
  • *Corresponding author for this work
  • School of Computer Science and Technology, Harbin Institute of Technology
  • Shandong University
  • Peking University
  • Guangdong Artificial Intelligence and Digital Economy Laboratory - Guangzhou

Research output: Contribution to journalArticlepeer-review

Abstract

Recently, deep clustering methods have achieved remarkable results compared to traditional clustering approaches. However, its performance remains constrained by the absence of annotations. A thought-provoking observation is that there is still a significant gap between deep clustering and semi-supervised classification methods. Even with only a few labeled samples, the accuracy of semi-supervised learning is much higher than that of clustering. Given that we can annotate a small number of samples in a certain unsupervised way, the clustering task can be naturally transformed into a semi-supervised setting, thereby achieving comparable performance. Based on this intuition, we propose ClusMatch, a unified positive and negative pseudo-label learning based semi-supervised learning framework, which is pluggable and can be applied to existing deep clustering methods. Specifically, we first leverage the pre-trained deep clustering network to compute predictions for all samples, and then design specialized selection strategies to pick out a few high-quality samples as labeled samples for supervised learning. For the unselected samples, the novel unified positive and negative pseudo-label learning is introduced to provide additional supervised signals for semi-supervised fine-tuning. We also propose an adaptive positive-negative threshold learning strategy to further enhance the confidence of generated pseudo-labels. Extensive experiments on six widely-used datasets and one large-scale dataset demonstrate the superiority of our proposed ClusMatch. For example, ClusMatch achieves a significant accuracy improvement of 5.4% over the state-of-the-art method ProPos on an average of these six datasets.

Original languageEnglish
Pages (from-to)9688-9701
Number of pages14
JournalIEEE Transactions on Pattern Analysis and Machine Intelligence
Volume47
Issue number11
DOIs
StatePublished - 2025
Externally publishedYes

Keywords

  • Deep clustering
  • positive and negative learning
  • semi-supervised learning

Fingerprint

Dive into the research topics of 'ClusMatch: Improving Deep Clustering by Unified Positive and Negative Pseudo-Label Learning'. Together they form a unique fingerprint.

Cite this