Skip to main navigation Skip to search Skip to main content

Image-text semantic learning for unsupervised cross-resolution person re-identification

  • Fuqi Liu
  • , Zhiqi Pang
  • , Chunyu Wang*
  • *Corresponding author for this work
  • Nanyang Technological University
  • Faculty of Computing, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Cross-resolution person re-identification (CR-ReID) focuses on matching person images of the same identity across different resolutions. Most existing CR-ReID methods rely on manually annotated identity labels for training. Although some researchers have proposed unsupervised CR-ReID (UCR-ReID) methods, the feature fusion techniques they rely on still require a large number of parameters and significant computational resources, limiting the widespread application of UCR-ReID technology. To address the aforementioned issues, we propose an image-text semantic learning (ITSL) method, which incorporates text semantics to enhance recognition performance. During the testing phase, ITSL requires only a single encoder to obtain resolution-invariant features. Specifically, ITSL first learns text features based on a visual-language model, and then utilizes the dual semantic matching module to match inter-resolution positive clusters in both the image and text modalities. During the optimization process, ITSL not only incorporates image semantic contrastive loss to facilitate cross-resolution alignment but also integrates text semantic contrastive loss to leverage text semantics for promoting resolution-invariance learning. Additionally, we design random region downsampling in ITSL, which further enhances the model's robustness to resolution gaps through data augmentation. Experimental results on multiple cross-resolution datasets show that ITSL not only outperforms existing unsupervised methods while maintaining efficiency, but also approaches the performance of earlier supervised methods on certain datasets.

Original languageEnglish
Article number128092
JournalExpert Systems with Applications
Volume286
DOIs
StatePublished - 15 Aug 2025
Externally publishedYes

Keywords

  • Person re-identification
  • Resolution gaps
  • Unsupervised learning
  • Vision-language models

Fingerprint

Dive into the research topics of 'Image-text semantic learning for unsupervised cross-resolution person re-identification'. Together they form a unique fingerprint.

Cite this