Skip to main navigation Skip to search Skip to main content

Health-oriented Multimodal Food Question Answering with Implicit and Explicit Knowledge

  • Menghao Hu
  • , Yaguang Song
  • , Xiaoshan Yang*
  • , Yaowei Wang
  • , Changsheng Xu
  • *Corresponding author for this work
  • Xinjiang University
  • Peng Cheng Laboratory
  • CAS - Institute of Automation
  • University of Chinese Academy of Sciences
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalArticlepeer-review

Abstract

Health-oriented food analysis has become a research hotspot in recent years because it can help people keep away from unhealthy diets. Remarkable advancements have been made in recipe retrieval, food recommendation, nutrition analysis, and calorie estimation. However, existing works still cannot well balance the individual preference and the health. Multimodal food question and answering (MFQA) presents substantial promise for practical applications, yet it remains underexplored. In this article, we introduce a health-oriented MFQA dataset with 9,000 Chinese question-answer pairs based on a multimodal food knowledge graph (MFKG) collected from a food-sharing Web site. Additionally, we propose a novel framework for MFQA in the health domain that leverages implicit general knowledge and explicit domain-specific knowledge. The framework comprises four key components: implicit general knowledge injection module (IGKIM), explicit domain-specific knowledge retrieval module (EDKRM), ranking module, and answer module. The IGKIM facilitates knowledge acquisition at both the feature and text levels. The EDKRM retrieves the most relevant candidate knowledge from the knowledge graph based on the given question. The ranking module sorts the results retrieved by EDKRM and further retrieve candidate knowledge relevant to the problem. Subsequently, the answer module thoroughly analyzes the multimodal information in the query along with the retrieved relevant knowledge to predict accurate answers. Extensive experimental results on the MFQA dataset demonstrate the effectiveness of our proposed method. The code and dataset are available at https://github.com/Wjianghai/HMFQA.

Original languageEnglish
Article number334
JournalACM Transactions on Multimedia Computing, Communications and Applications
Volume21
Issue number12
DOIs
StatePublished - 22 Nov 2025
Externally publishedYes

Keywords

  • Food Analysis
  • Health
  • Knowledge Representation

Fingerprint

Dive into the research topics of 'Health-oriented Multimodal Food Question Answering with Implicit and Explicit Knowledge'. Together they form a unique fingerprint.

Cite this