Skip to main navigation Skip to search Skip to main content

Query-based Collaborative Multimodal Token Pruning for Audio-Visual Question Answering

  • Xilin Qin
  • , Dexiang Hong
  • , Weidong Chen*
  • , Cheng Ye
  • , Xinyan Liu
  • , Peipei Song
  • , Lei Zhang
  • *Corresponding author for this work
  • University of Science and Technology of China
  • School of Computer Science and Technology (School of Software), Harbin Institute of Technology Weihai

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Audio-Visual Question Answering (AVQA) requires complex reasoning across auditory and visual modalities. While recent advancements leverage sophisticated spatio-temporal representations, cross-modal fusion, and question-aware mechanisms to boost performance, they frequently neglect the computational burden and significant information redundancy inherent in processing long, dense video sequences. This inefficiency stems from frame-by-frame processing where critical information is often sparse. To bridge this efficiency gap, this paper introduces a novel Query-based Collaborative Multimodal Token Pruning strategy. Our method exploits the coupling relationship across modalities between layers to collaboratively assesses multimodal token importance required for the current state and update the query accordingly by the proposed Adaptive Multimodal Token Pruning. This query-based collaboration enables highly efficient, focused reasoning. Experiments on the MUSIC-AVQA dataset demonstrate that our method prunes up to 80% of multimodal tokens without accuracy degradation, validating our query-based collaborative multimodal token pruning as a highly effective strategy for building more efficient and robust AVQA models.

Original languageEnglish
Title of host publication2025 4th International Conference on Artificial Intelligence, Human-Computer Interaction and Robotics, AIHCIR 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331554934
DOIs
StatePublished - 2025
Externally publishedYes
Event2025 4th International Conference on Artificial Intelligence, Human-Computer Interaction and Robotics, AIHCIR 2025 - Ningbo, China
Duration: 28 Nov 202530 Nov 2025

Publication series

Name2025 4th International Conference on Artificial Intelligence, Human-Computer Interaction and Robotics, AIHCIR 2025

Conference

Conference2025 4th International Conference on Artificial Intelligence, Human-Computer Interaction and Robotics, AIHCIR 2025
Country/TerritoryChina
CityNingbo
Period28/11/2530/11/25

Keywords

  • Adaptive Multimodal Token Pruning
  • Audio-Visual Question Answering

Fingerprint

Dive into the research topics of 'Query-based Collaborative Multimodal Token Pruning for Audio-Visual Question Answering'. Together they form a unique fingerprint.

Cite this