Skip to main navigation Skip to search Skip to main content

Semi-supervised Visual Feature Integration for Language Models through Sentence Visualization

  • Harbin Institute of Technology Shenzhen
  • Peng Cheng Laboratory

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Integrating visual features has been proved useful for natural language understanding tasks. Nevertheless, most existing multimodal language models highly rely on training on aligned image and text data. In this paper, we propose a novel semi-supervised visual integration framework for pre-trained language models. In the framework, the visual features are obtained through a sentence visualization and vision-language fusion mechanism. The uniqueness includes: 1) the integration is conducted via a semi-supervised framework and does not require aligned images for the processed sentences. 2) the framework works as an auxiliary component, and will not affect the language processing ability of the integrated language model. Experimental results on both natural language inference and reading comprehension tasks demonstrate that our framework improves the strong baseline language models. Considering that our framework only requires an image database, and does not require aligned images for the processed texts, it provides a feasible way for multimodal language learning.

Original languageEnglish
Title of host publicationICMI 2021 - Proceedings of the 2021 International Conference on Multimodal Interaction
PublisherAssociation for Computing Machinery, Inc
Pages682-686
Number of pages5
ISBN (Electronic)9781450384810
DOIs
StatePublished - 18 Oct 2021
Externally publishedYes
Event23rd ACM International Conference on Multimodal Interaction, ICMI 2021 - Virtual, Online, Canada
Duration: 18 Oct 202122 Oct 2021

Publication series

NameICMI 2021 - Proceedings of the 2021 International Conference on Multimodal Interaction

Conference

Conference23rd ACM International Conference on Multimodal Interaction, ICMI 2021
Country/TerritoryCanada
CityVirtual, Online
Period18/10/2122/10/21

Keywords

  • language model
  • multimodal fusion
  • vision and language

Fingerprint

Dive into the research topics of 'Semi-supervised Visual Feature Integration for Language Models through Sentence Visualization'. Together they form a unique fingerprint.

Cite this