Skip to main navigation Skip to search Skip to main content

Exploiting Multimodal Knowledge Graph for Multimodal Machine Translation

  • Tianjiao Xu
  • , Xuebo Liu*
  • , Derek F. Wong*
  • , Yue Zhang
  • , Lidia S. Chao
  • , Min Zhang
  • , Tian Gan
  • *Corresponding author for this work
  • Shandong University
  • Harbin Institute of Technology Shenzhen
  • University of Macau
  • Westlake University

Research output: Contribution to journalArticlepeer-review

Abstract

A neural Multimodal Machine Translation (MMT) system utilizes multimodal information, particularly images, to enhance traditional text-only models and achieve superior performance. However, the effectiveness of MMT heavily depends on the availability of extensive collections of bilingual parallel sentence pairs and manually annotated images, which poses a challenge due to the scarcity of such pairs. To address this issue, we propose incorporating the Multimodal Knowledge Graph (MMKG) for data augmentation in MMT. By utilizing MMKG as an additional source of knowledge, we can overcome the limitations of existing sentence-image pairings. This allows us to expand the original parallel corpus and generate corresponding images, creating new synthetic data pairs that facilitate effective data augmentation. Experiments conducted on two translation datasets, Multi30k and IKEA, demonstrate that the proposed MMKG enhancement method significantly improves performance across multiple baseline methods, ultimately outperforming all baseline approaches. Additionally, experiments under low-resource conditions reveal that our method achieves exceptional enhancement effects in low-resource corpora, surpassing other data augmentation baseline methods. These results indicate the efficacy and potential of the proposed method for enhancing the performance of multimodal models across diverse datasets.

Original languageEnglish
Pages (from-to)1677-1688
Number of pages12
JournalIEEE Transactions on Multimedia
Volume28
DOIs
StatePublished - 2026
Externally publishedYes

Keywords

  • Multimodal machine translation
  • data augmentation
  • low-resource
  • multimodal knowledge graph

Fingerprint

Dive into the research topics of 'Exploiting Multimodal Knowledge Graph for Multimodal Machine Translation'. Together they form a unique fingerprint.

Cite this