Skip to main navigation Skip to search Skip to main content

Cross-modal recipe retrieval via parallel- and cross-attention networks learning

  • Da Cao
  • , Jingjing Chu
  • , Ningbo Zhu*
  • , Liqiang Nie
  • *Corresponding author for this work
  • Hunan University
  • Shandong University

Research output: Contribution to journalArticlepeer-review

Abstract

Cross-modal recipe retrieval refers to the problem of retrieving a food image from a list of image candidates given a textual recipe as the query, or the reverse side. However, existing cross-modal recipe retrieval approaches mostly focus on learning the representations of images and recipes independently and sewing them up by projecting them into a common space. Such methods overlook the interplay between images and recipes, resulting in the suboptimal retrieval performance. Toward this end, we study the problem of cross-modal recipe retrieval from the viewpoint of parallel- and cross-attention networks learning. Specifically, we first exploit a parallel-attention network to independently learn the attention weights of components in images and recipes. Thereafter, a cross-attention network is proposed to explicitly learn the interplay between images and recipes, which simultaneously considers word-guided image attention and image-guided word attention. Lastly, the learnt representations of images and recipes stemming from parallel- and cross-attention networks are elaborately connected and optimized using a pairwise ranking loss. By experimenting on two datasets, we demonstrate the effectiveness and rationality of our proposed solution on the scope of both overall performance comparison and micro-level analyses.

Original languageEnglish
Article number105428
JournalKnowledge-Based Systems
Volume193
DOIs
StatePublished - 6 Apr 2020
Externally publishedYes

Keywords

  • Cross-attention network
  • Cross-modal retrieval
  • Parallel-attention network
  • Recipe retrieval

Fingerprint

Dive into the research topics of 'Cross-modal recipe retrieval via parallel- and cross-attention networks learning'. Together they form a unique fingerprint.

Cite this