Skip to main navigation Skip to search Skip to main content

Synergistic Prompting Learning for Human-Object Interaction Detection

  • Jinguo Luo
  • , Weihong Ren*
  • , Zhiyong Wang
  • , Xi'ai Chen
  • , Huijie Fan
  • , Zhi Han
  • , Honghai Liu
  • *Corresponding author for this work
  • School of Biomedical Engineering, Harbin Institute of Technology Shenzhen
  • CAS - Shenyang Institute of Automation

Research output: Contribution to journalArticlepeer-review

Abstract

Human-Object Interaction (HOI) detection, as a foundational task in human-centric understanding, aims to detect interactive triplets in real-world scenarios. To better distinguish diverse HOIs within an open-world context, current HOI detectors utilize pre-trained Visual-Language Models (VLMs) to extract prior knowledge through textual prompts (i.e., descriptive texts for each HOI instance). However, relying on predetermined descriptive texts, such approaches only acquire a fixed set of textual knowledge for HOI prediction, consequently resulting in inferior performance and limited generalization. To remedy this, we propose a novel VLM-based method, which jointly performs prompting learning from both visual and textual perspectives and synergizes visual-textual prompting for HOI detection. Initially, we design a hierarchical adaptation architecture to perform progressive prompting: visual prompting is facilitated through gradual token migration from VLM’s image encoder, while textual prompting is initialized with progressively leveled interaction descriptions. In addition, to synergize the visual-textual prompting learning, a text-supervising and image-tuning loop is introduced, in which the text-supervising stage guides visual prompting learning through contrastive learning and the image-tuning stage refines textual prompting by modal matching. Finally, we employ an interaction-aware knowledge merging mechanism to effectively transfer visual-textual knowledge encapsulated within synergistic prompting for HOI detection. Extensive experiments on two benchmarks demonstrate that our proposed method outperforms the state-of-the-art ones, under both supervised and zero-shot settings.

Original languageEnglish
Pages (from-to)5710-5724
Number of pages15
JournalIEEE Transactions on Image Processing
Volume34
DOIs
StatePublished - 2025
Externally publishedYes

Keywords

  • Human–object interaction
  • prior knowledge
  • prompting learning
  • visual-textual synergy

Fingerprint

Dive into the research topics of 'Synergistic Prompting Learning for Human-Object Interaction Detection'. Together they form a unique fingerprint.

Cite this