Skip to main navigation Skip to search Skip to main content

Modular Graph Attention Network for Complex Visual Relational Reasoning

  • Yihan Zheng
  • , Zhiquan Wen
  • , Mingkui Tan
  • , Runhao Zeng
  • , Qi Chen
  • , Yaowei Wang*
  • , Qi Wu
  • *Corresponding author for this work
  • South China University of Technology
  • Pengcheng Laboratory
  • Adelaide University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Visual Relational Reasoning is crucial for many vision-and-language based tasks, such as Visual Question Answering and Vision Language Navigation. In this paper, we consider reasoning on complex referring expression comprehension (c-REF) task that seeks to localise the target objects in an image guided by complex queries. Such queries often contain complex logic and thus impose two key challenges for reasoning: (i) It can be very difficult to comprehend the query since it often refers to multiple objects and describes complex relationships among them. (ii) It is non-trivial to reason among multiple objects guided by the query and localise the target correctly. To address these challenges, we propose a novel Modular Graph Attention Network (MGA-Net). Specifically, to comprehend the long queries, we devise a language attention network to decompose them into four types: basic attributes, absolute location, visual relationship and relative locations, which mimics the human language understanding mechanism. Moreover, to capture the complex logic in a query, we construct a relational graph to represent the visual objects and their relationships, and propose a multi-step reasoning method to progressively understand the complex logic. Extensive experiments on CLEVR-Ref+, GQA and CLEVR-CoGenT datasets demonstrate the superior reasoning performance of our MGA-Net.

Original languageEnglish
Title of host publicationComputer Vision – ACCV 2020 - 15th Asian Conference on Computer Vision, 2020, Revised Selected Papers
EditorsHiroshi Ishikawa, Cheng-Lin Liu, Tomas Pajdla, Jianbo Shi
PublisherSpringer Science and Business Media Deutschland GmbH
Pages137-153
Number of pages17
ISBN (Print)9783030695439
DOIs
StatePublished - 2021
Externally publishedYes
Event15th Asian Conference on Computer Vision, ACCV 2020 - Virtual, Online, Japan
Duration: 30 Nov 20204 Dec 2020

Publication series

NameLecture Notes in Computer Science
Volume12627 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference15th Asian Conference on Computer Vision, ACCV 2020
Country/TerritoryJapan
CityVirtual, Online
Period30/11/204/12/20

Fingerprint

Dive into the research topics of 'Modular Graph Attention Network for Complex Visual Relational Reasoning'. Together they form a unique fingerprint.

Cite this