TY - GEN
T1 - Query-based Collaborative Multimodal Token Pruning for Audio-Visual Question Answering
AU - Qin, Xilin
AU - Hong, Dexiang
AU - Chen, Weidong
AU - Ye, Cheng
AU - Liu, Xinyan
AU - Song, Peipei
AU - Zhang, Lei
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Audio-Visual Question Answering (AVQA) requires complex reasoning across auditory and visual modalities. While recent advancements leverage sophisticated spatio-temporal representations, cross-modal fusion, and question-aware mechanisms to boost performance, they frequently neglect the computational burden and significant information redundancy inherent in processing long, dense video sequences. This inefficiency stems from frame-by-frame processing where critical information is often sparse. To bridge this efficiency gap, this paper introduces a novel Query-based Collaborative Multimodal Token Pruning strategy. Our method exploits the coupling relationship across modalities between layers to collaboratively assesses multimodal token importance required for the current state and update the query accordingly by the proposed Adaptive Multimodal Token Pruning. This query-based collaboration enables highly efficient, focused reasoning. Experiments on the MUSIC-AVQA dataset demonstrate that our method prunes up to 80% of multimodal tokens without accuracy degradation, validating our query-based collaborative multimodal token pruning as a highly effective strategy for building more efficient and robust AVQA models.
AB - Audio-Visual Question Answering (AVQA) requires complex reasoning across auditory and visual modalities. While recent advancements leverage sophisticated spatio-temporal representations, cross-modal fusion, and question-aware mechanisms to boost performance, they frequently neglect the computational burden and significant information redundancy inherent in processing long, dense video sequences. This inefficiency stems from frame-by-frame processing where critical information is often sparse. To bridge this efficiency gap, this paper introduces a novel Query-based Collaborative Multimodal Token Pruning strategy. Our method exploits the coupling relationship across modalities between layers to collaboratively assesses multimodal token importance required for the current state and update the query accordingly by the proposed Adaptive Multimodal Token Pruning. This query-based collaboration enables highly efficient, focused reasoning. Experiments on the MUSIC-AVQA dataset demonstrate that our method prunes up to 80% of multimodal tokens without accuracy degradation, validating our query-based collaborative multimodal token pruning as a highly effective strategy for building more efficient and robust AVQA models.
KW - Adaptive Multimodal Token Pruning
KW - Audio-Visual Question Answering
UR - https://www.scopus.com/pages/publications/105035992006
U2 - 10.1109/AIHCIR67580.2025.11405267
DO - 10.1109/AIHCIR67580.2025.11405267
M3 - 会议稿件
AN - SCOPUS:105035992006
T3 - 2025 4th International Conference on Artificial Intelligence, Human-Computer Interaction and Robotics, AIHCIR 2025
BT - 2025 4th International Conference on Artificial Intelligence, Human-Computer Interaction and Robotics, AIHCIR 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 4th International Conference on Artificial Intelligence, Human-Computer Interaction and Robotics, AIHCIR 2025
Y2 - 28 November 2025 through 30 November 2025
ER -