TY - GEN
T1 - Universal Scene Graph Generation via Semantic Feature Alignment
AU - Zhang, Xiangyu
AU - Qiu, Guoxi
AU - Xu, Yong
AU - Wang, Jinghua
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Scene Graph Generation (SGG) offers a valuable structured representation for various computer vision applications. With advancements in visual and textual alignment, researchers are exploring both open-vocabulary object detection to address unseen object categories and open-vocabulary relation detection to manage unseen relation categories. To address real-world applications, we propose a new universal scene graph generation framework (U-SGG) that handles both unseen object categories and unseen relation categories. Our method detects unseen objects by leveraging the aligned representations of images and text. To effectively address unseen relations, we build the relation prediction head using a Visual Language Transformer, which takes into consideration not only the visual and textual features, but also the attribute and spatial features of the objects. To alleviate the problem of insufficient quantity of training data in SGG tasks, we propose to parse image captions into relation triples, thereby enriching the categories of both object and relation. Using image caption data with diverse object and relation categories, we successfully achieve universal scene graph generation. Comprehensive experimental results on the Visual Genome benchmark demonstrate the effectiveness and superiority of the proposed method.
AB - Scene Graph Generation (SGG) offers a valuable structured representation for various computer vision applications. With advancements in visual and textual alignment, researchers are exploring both open-vocabulary object detection to address unseen object categories and open-vocabulary relation detection to manage unseen relation categories. To address real-world applications, we propose a new universal scene graph generation framework (U-SGG) that handles both unseen object categories and unseen relation categories. Our method detects unseen objects by leveraging the aligned representations of images and text. To effectively address unseen relations, we build the relation prediction head using a Visual Language Transformer, which takes into consideration not only the visual and textual features, but also the attribute and spatial features of the objects. To alleviate the problem of insufficient quantity of training data in SGG tasks, we propose to parse image captions into relation triples, thereby enriching the categories of both object and relation. Using image caption data with diverse object and relation categories, we successfully achieve universal scene graph generation. Comprehensive experimental results on the Visual Genome benchmark demonstrate the effectiveness and superiority of the proposed method.
KW - Multimodal learning and alignment
KW - Scene graph generation
UR - https://www.scopus.com/pages/publications/105022616859
U2 - 10.1109/ICME59968.2025.11209357
DO - 10.1109/ICME59968.2025.11209357
M3 - 会议稿件
AN - SCOPUS:105022616859
T3 - Proceedings - IEEE International Conference on Multimedia and Expo
BT - 2025 IEEE International Conference on Multimedia and Expo
PB - IEEE Computer Society
T2 - 2025 IEEE International Conference on Multimedia and Expo, ICME 2025
Y2 - 30 June 2025 through 4 July 2025
ER -