Abstract
Text-to-motion retrieval (TMR) matches human motions and textual descriptions by learning larger similarity scores for positive pairs and smaller similarity scores for negative pairs, yet remains challenging due to inherent cross-modal discrepancies. Inspired by the observation that human movements are highly correlated with the corresponding descriptions in the text-to-motion generation (TMG) task, this paper presents a novel generation-assisted TMR framework that leverages this intrinsic text-motion correlation in TMG. Our key innovation is RetNet, a retrieval model that utilizes the encoder from a diffusion-based generative UNet to directly learn similarity scores from high-level semantic features. The framework jointly optimizes two branches: (1) GenNet, a UNet comprising Transformer encoder–decoder blocks for text-to-motion synthesis, and (2) RetNet, which shares GenNet's encoder and employs an MLP head to regress high-level text-motion features to similarity scores. In encoder layers of RetNet, positive text-motion pairs exhibit stronger feature activation, enabling higher predicted similarity. Furthermore, to enhance the generation quality of GenNet, we employ a frozen RetNet to maximize text-motion mutual information between conditional texts and generated motions for better alignment. For efficient deployment, we distill the knowledge of RetNet into an existing fast retrieval framework, called Dual Encoder Network (DENet), through hybrid supervision of hard labels and RetNet's soft similarity scores. Experimental results on HumanML3D and KIT-ML datasets demonstrate that our method achieves the state-of-the-art (SOTA) performance using only the global motion features.
| Original language | English |
|---|---|
| Article number | 113983 |
| Journal | Pattern Recognition |
| Volume | 180 |
| DOIs | |
| State | Published - Dec 2026 |
| Externally published | Yes |
Keywords
- Diffusion models
- Text-to-motion generation
- Text-to-motion retrieval
Fingerprint
Dive into the research topics of 'Text-to-motion retrieval by text-to-motion generation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver