Skip to main navigation Skip to search Skip to main content

Digging out Discrimination Information from Generated Samples for Robust Visual Question Answering

  • Zhiquan Wen
  • , Yaowei Wang*
  • , Mingkui Tan*
  • , Qingyao Wu
  • , Qi Wu
  • *Corresponding author for this work
  • South China University of Technology
  • PengCheng Laboratory
  • University of Adelaide

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Visual Question Answering (VQA) aims to answer a textual question based on a given image. Nevertheless, recent studies have shown that VQA models tend to capture the biases to answer the question, instead of using the reasoning ability, resulting in poor generalisation ability. To alleviate the issue, some existing methods consider the natural distribution of the data, and construct samples to balance the dataset, achieving remarkable performance. However, these methods may encounter some limitations: 1) rely on additional annotations, 2) the generated samples may be inaccurate, e.g., assigned wrong answers, and 3) ignore the power of positive samples. In this paper, we propose a method to Dig out Discrimination information from Generated samples (DDG) to address the above limitations. Specifically, we first construct positive and negative samples in vision and language modalities, without using additional annotations. Then, we introduce a knowledge distillation mechanism to promote the learning of the original samples by the positive samples. Moreover, we impel the VQA models to focus on vision and language modalities using the negative samples. Experimental results on the VQA-CP v2 and VQA v2 datasets show the effectiveness of our DDG.

Original languageEnglish
Title of host publicationFindings of the Association for Computational Linguistics, ACL 2023
PublisherAssociation for Computational Linguistics (ACL)
Pages6910-6928
Number of pages19
ISBN (Electronic)9781959429623
DOIs
StatePublished - 2023
Externally publishedYes
EventFindings of the Association for Computational Linguistics, ACL 2023 - Toronto, Canada
Duration: 9 Jul 202314 Jul 2023

Publication series

NameProceedings of the Annual Meeting of the Association for Computational Linguistics
ISSN (Print)0736-587X

Conference

ConferenceFindings of the Association for Computational Linguistics, ACL 2023
Country/TerritoryCanada
CityToronto
Period9/07/2314/07/23

Fingerprint

Dive into the research topics of 'Digging out Discrimination Information from Generated Samples for Robust Visual Question Answering'. Together they form a unique fingerprint.

Cite this