Skip to main navigation Skip to search Skip to main content

Universal Scene Graph Generation via Semantic Feature Alignment

  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Scene Graph Generation (SGG) offers a valuable structured representation for various computer vision applications. With advancements in visual and textual alignment, researchers are exploring both open-vocabulary object detection to address unseen object categories and open-vocabulary relation detection to manage unseen relation categories. To address real-world applications, we propose a new universal scene graph generation framework (U-SGG) that handles both unseen object categories and unseen relation categories. Our method detects unseen objects by leveraging the aligned representations of images and text. To effectively address unseen relations, we build the relation prediction head using a Visual Language Transformer, which takes into consideration not only the visual and textual features, but also the attribute and spatial features of the objects. To alleviate the problem of insufficient quantity of training data in SGG tasks, we propose to parse image captions into relation triples, thereby enriching the categories of both object and relation. Using image caption data with diverse object and relation categories, we successfully achieve universal scene graph generation. Comprehensive experimental results on the Visual Genome benchmark demonstrate the effectiveness and superiority of the proposed method.

Original languageEnglish
Title of host publication2025 IEEE International Conference on Multimedia and Expo
Subtitle of host publicationJourney to the Center of Machine Imagination, ICME 2025 - Conference Proceedings
PublisherIEEE Computer Society
ISBN (Electronic)9798331594954
DOIs
StatePublished - 2025
Externally publishedYes
Event2025 IEEE International Conference on Multimedia and Expo, ICME 2025 - Nantes, France
Duration: 30 Jun 20254 Jul 2025

Publication series

NameProceedings - IEEE International Conference on Multimedia and Expo
ISSN (Print)1945-7871
ISSN (Electronic)1945-788X

Conference

Conference2025 IEEE International Conference on Multimedia and Expo, ICME 2025
Country/TerritoryFrance
CityNantes
Period30/06/254/07/25

Keywords

  • Multimodal learning and alignment
  • Scene graph generation

Fingerprint

Dive into the research topics of 'Universal Scene Graph Generation via Semantic Feature Alignment'. Together they form a unique fingerprint.

Cite this