Skip to main navigation Skip to search Skip to main content

Deep Multi-Modal Hashing With Semantic Enhancement for Multi-Label Micro-Video Retrieval

  • Peiguang Jing
  • , Haoyi Sun
  • , Liqiang Nie
  • , Yun Li*
  • , Yuting Su
  • *Corresponding author for this work
  • Tianjin University
  • School of Computer Science and Technology, Harbin Institute of Technology
  • Guangxi University of Finance and Economics
  • Guangxi Key Laboratory of Big Data in Finance and Economics

Research output: Contribution to journalArticlepeer-review

Abstract

The pressing need for low storage and high efficiency has significantly propelled the advancement of deep hashing techniques in the realm of large-scale search and retrieval tasks. As one of the most prevailing forms of user-generated contents, micro-videos usually represent more complicated multi-modal behaviors that are further challenged in multi-label retrieval. Existing multi-modal hashing methods tend to prioritize the complementarity and consistency in multi-modal fusion, while neglecting the completeness problem. In this paper, we propose a deep multi-modal hashing with semantic enhancement (DMHSE) method that effectively integrates complete multi-modal representation learning with discriminative binary coding by means of collaboration between two distinct encoders, FoldCoder and HashCoder. FoldCoder translates latent multi-modal representation learning to a degradation process through mimicking data transmitting. Further, it incorporates a prompt learning paradigm to maximize the utilization of multi-label semantics for guiding representation learning. HashCoder combines pairwise and central constraints to ensure more discriminative hashing results. Pairwise constraint preserves the original local relevance structure, while central constraint tackles the problem of semantic ambiguity in multi-label data by leveraging the global label distribution. Experimental results demonstrate that DMHSE achieves superior performance in multi-label micro-video retrieval tasks.

Original languageEnglish
Pages (from-to)5080-5091
Number of pages12
JournalIEEE Transactions on Knowledge and Data Engineering
Volume36
Issue number10
DOIs
StatePublished - 2024
Externally publishedYes

Keywords

  • Deep hashing
  • micro-video retrieval
  • multi-label
  • multi-modality

Fingerprint

Dive into the research topics of 'Deep Multi-Modal Hashing With Semantic Enhancement for Multi-Label Micro-Video Retrieval'. Together they form a unique fingerprint.

Cite this