Skip to main navigation Skip to search Skip to main content

Improving multi-instance learning with hierarchical attention and frequency-domain hard sample distillation

  • Ting Xiao
  • , Minqian Sun
  • , Yiqing Xia
  • , Hai Yang
  • , Zhe Wang*
  • , Peng Liu
  • *Corresponding author for this work
  • East China University of Science and Technology
  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Weakly supervised multiple instance learning (MIL) has gained widespread employment in whole slide image (WSI) classification, with most methods favoring the bag-based approach due to its good performance. Most approaches incorporate an attention mechanism to generate bag-level representations. However, due to a significant imbalance between the number of positive and negative instances within positive bags, attention-based MIL methods still have a cascade of issues: (1) highly imbalanced instances lead to attention over-concentration; (2) attention over-concentration leads to difficulty in balancing salient and hard instances, and results in misguided attention to irrelevant patterns. This paper proposes a novel MIL framework, named HAFD-MIL, which incorporates hierarchical attention and frequency-domain hard sample distillation. This framework addresses these challenges by jointly considering sample categories and sample difficulty. Specifically, from the perspective of sample categories, we first classify instances within each pseudo-bag as “trend-to-positive”, “trend-to-negative”, or “weak negative” based on negative instance prototypes clustered from negative bags. Then we propose a graded attention distillation module to process the above-classified instances separately to reduce the influence of negative instances on positive instances, along with a novel dynamic label attention loss to prevent attention over-concentration. From the perspective of sample difficulty, we propose a spectrum attention distillation module designed to extract information from hard samples and a redundant feature filter module to minimize interference from irrelevant information typically involved in the attention mechanism. Experimental results on two WSI datasets with three pre-trained backbones demonstrate the superiority of our HAFD-MIL framework and exhibit a wider range of attention areas.

Original languageEnglish
Article number116286
JournalKnowledge-Based Systems
Volume347
DOIs
StatePublished - 19 Jul 2026
Externally publishedYes

Keywords

  • Attention regularization
  • Hard sample mining
  • Histological whole slide image
  • Multiple instance learning

Fingerprint

Dive into the research topics of 'Improving multi-instance learning with hierarchical attention and frequency-domain hard sample distillation'. Together they form a unique fingerprint.

Cite this