Skip to main navigation Skip to search Skip to main content

Text-to-motion retrieval by text-to-motion generation

  • Honghu Pan
  • , Qianqian Wang
  • , Guoqing Zhu
  • , Yongyong Chen*
  • *Corresponding author for this work
  • Hunan University
  • Northeastern University China
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Text-to-motion retrieval (TMR) matches human motions and textual descriptions by learning larger similarity scores for positive pairs and smaller similarity scores for negative pairs, yet remains challenging due to inherent cross-modal discrepancies. Inspired by the observation that human movements are highly correlated with the corresponding descriptions in the text-to-motion generation (TMG) task, this paper presents a novel generation-assisted TMR framework that leverages this intrinsic text-motion correlation in TMG. Our key innovation is RetNet, a retrieval model that utilizes the encoder from a diffusion-based generative UNet to directly learn similarity scores from high-level semantic features. The framework jointly optimizes two branches: (1) GenNet, a UNet comprising Transformer encoder–decoder blocks for text-to-motion synthesis, and (2) RetNet, which shares GenNet's encoder and employs an MLP head to regress high-level text-motion features to similarity scores. In encoder layers of RetNet, positive text-motion pairs exhibit stronger feature activation, enabling higher predicted similarity. Furthermore, to enhance the generation quality of GenNet, we employ a frozen RetNet to maximize text-motion mutual information between conditional texts and generated motions for better alignment. For efficient deployment, we distill the knowledge of RetNet into an existing fast retrieval framework, called Dual Encoder Network (DENet), through hybrid supervision of hard labels and RetNet's soft similarity scores. Experimental results on HumanML3D and KIT-ML datasets demonstrate that our method achieves the state-of-the-art (SOTA) performance using only the global motion features.

Original languageEnglish
Article number113983
JournalPattern Recognition
Volume180
DOIs
StatePublished - Dec 2026
Externally publishedYes

Keywords

  • Diffusion models
  • Text-to-motion generation
  • Text-to-motion retrieval

Fingerprint

Dive into the research topics of 'Text-to-motion retrieval by text-to-motion generation'. Together they form a unique fingerprint.

Cite this