Abstract
Effective retrieval and structuring of heterogeneous data have grown more difficult due to the exponential development of multimedia data. The surge in data volume emphasizes the importance of efficient cross-modal hashing techniques, known for their rapid retrieval speed and minimal storage requirements, which have garnered attention recently. However, existing unsupervised cross-modal hashing methods often fail to capture latent semantic structures and meaningful modality interactions, which limits their retrieval performance. To address these challenges, we propose Attention-driven Contrastive Learning for Cross-Modal Hashing via Prototypical Separation (ACoPSe). The method introduces a modality-aware fusion mechanism to enhance cross-modal feature interaction and a prototype alignment strategy that reduces heterogeneity at the cluster level by leveraging pseudo-labels derived from clustering. Extensive experiments demonstrate that our method achieves comparable performance to state-of-the-art approaches.
| Original language | English |
|---|---|
| Article number | 104078 |
| Journal | Information Fusion |
| Volume | 130 |
| DOIs | |
| State | Published - Jun 2026 |
| Externally published | Yes |
Keywords
- Common Hamming space
- Cross-modal retrieval
- Pseudo-label
- Unsupervised learning
Fingerprint
Dive into the research topics of 'Attention-driven contrastive learning for cross-modal hashing with prototypical separation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver