Skip to main navigation Skip to search Skip to main content

Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image

  • Yu Zhao
  • , Hao Fei*
  • , Xiangtai Li
  • , Libo Qin
  • , Jiayi Ji
  • , Hongyuan Zhu
  • , Meishan Zhang
  • , Min Zhang
  • , Jianguo Wei
  • *Corresponding author for this work
  • Tianjin University
  • National University of Singapore
  • ByteDance Ltd.
  • Central South University
  • Agency for Science, Technology and Research, Singapore
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalConference articlepeer-review

Abstract

In the visual spatial understanding (VSU) area, spatial image-to-text (SI2T) and spatial text-to-image (ST2I) are two fundamental tasks that appear in dual form. Existing methods for standalone SI2T or ST2I perform imperfectly in spatial understanding, due to the difficulty of 3D-wise spatial feature modeling. In this work, we consider modeling the SI2T and ST2I together under a dual learning framework. During the dual framework, we then propose to represent the 3D spatial scene features with a novel 3D scene graph (3DSG) representation that can be shared and beneficial to both tasks. Further, inspired by the intuition that the easier 3D→image and 3D→text processes also exist symmetrically in the ST2I and SI2T, respectively, we propose the Spatial Dual Discrete Diffusion (SD3) framework, which utilizes the intermediate features of the 3D→X processes to guide the hard X→3D processes, such that the overall ST2I and SI2T will benefit each other. On the visual spatial understanding dataset VSD, our system outperforms the mainstream T2I and I2T methods significantly. Further in-depth analysis reveals how our dual learning strategy advances.

Original languageEnglish
JournalAdvances in Neural Information Processing Systems
Volume37
StatePublished - 2024
Externally publishedYes
Event38th Conference on Neural Information Processing Systems, NeurIPS 2024 - Vancouver, Canada
Duration: 9 Dec 202415 Dec 2024

Fingerprint

Dive into the research topics of 'Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image'. Together they form a unique fingerprint.

Cite this