Skip to main navigation Skip to search Skip to main content

Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning

  • Min Cao
  • , Xinyu Zhou
  • , Ding Jiang
  • , Bo Du
  • , Mang Ye*
  • , Min Zhang
  • *Corresponding author for this work
  • Soochow University
  • Wuhan University

Research output: Contribution to journalArticlepeer-review

Abstract

Text-to-image person retrieval (TIPR) aims to identify the target person using textual descriptions, facing challenge in modality heterogeneity. Prior works have attempted to address it by developing cross-modal global or local alignment strategies. However, global methods typically overlook fine-grained cross-modal differences, whereas local methods require prior information to explore explicit part alignments. Additionally, current methods are English-centric, restricting their application in multilingual contexts. To alleviate these issues, we pioneer a multilingual TIPR task by developing a multilingual TIPR benchmark, for which we leverage large language models for initial translations and refine them by integrating domain-specific knowledge. Correspondingly, we propose Bi-IRRA: a Bidirectional Implicit Relation Reasoning and Aligning framework to learn alignment across languages and modalities. Within Bi-IRRA, a bidirectional implicit relation reasoning module enables bidirectional prediction of masked image and text, implicitly enhancing the modeling of local relations across languages and modalities, a multi-dimensional global alignment module is integrated to bridge the modality heterogeneity. The proposed method achieves new state-of-the-art results on all multilingual TIPR datasets.

Original languageEnglish
Pages (from-to)1961-1977
Number of pages17
JournalIEEE Transactions on Pattern Analysis and Machine Intelligence
Volume48
Issue number2
DOIs
StatePublished - Feb 2026
Externally publishedYes

Keywords

  • Text-to-image person retrieval
  • multilingual image-text learning
  • person re-identification

Fingerprint

Dive into the research topics of 'Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning'. Together they form a unique fingerprint.

Cite this