Skip to main navigation Skip to search Skip to main content

Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries

  • Yin Wu*
  • , Quanyu Long
  • , Jing Li
  • , Jianfei Yu
  • , Wenya Wang
  • *Corresponding author for this work
  • Nanyang Technological University
  • Harbin Institute of Technology Shenzhen
  • Nanjing University of Science and Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Retrieval-augmented generation (RAG) augments large language models (LLMs) with external knowledge to tackle knowledge-intensive question answering. While several benchmarks evaluate multimodal LLMs (MLLMs) under multimodal RAG settings, they predominantly retrieve from textual corpora and do not explicitly assess how models exploit visual evidence during answer generation. Consequently, there still lacks benchmark that cleanly isolates and measures the contribution of retrieved images in a visual knowledge-intensive RAG pipeline. We introduce Visual-RAG, a question-answering benchmark that targets visually-grounded, knowledge-intensive questions in a visual evidence-centric manner. Unlike prior work, Visual-RAG requires text-to-image retrieval and the integration of retrieved clue images whose pixel content explicitly encodes the visual knowledge necessary for answer generation. With Visual-RAG, we evaluate five open-source and three proprietary MLLMs and find that current systems still substantially underutilize the visual information available in retrieved images. Despite clear opportunities for multimodal evidence integration, state-of-the-art models struggle to extract and exploit fine-grained visual knowledge, and text-to-image retrieval itself remains challenging even under constrained entity-level corpora. These results underscore the need for improved visual retrieval, grounding, and attribution in multimodal RAG. Visual-RAG is publicly available at: github.com/visual-rag/visual-rag

Original languageEnglish
Title of host publicationSIGIR 2026 - Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval
PublisherAssociation for Computing Machinery, Inc
Pages3473-3480
Number of pages8
ISBN (Electronic)9798400725999
DOIs
StatePublished - 19 Jul 2026
Externally publishedYes
Event49th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2026 - Melbourne, Australia
Duration: 20 Jul 202624 Jul 2026

Publication series

NameSIGIR 2026 - Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval

Conference

Conference49th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2026
Country/TerritoryAustralia
CityMelbourne
Period20/07/2624/07/26

Keywords

  • multimodal question answering
  • retrieval augmented generation

Fingerprint

Dive into the research topics of 'Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries'. Together they form a unique fingerprint.

Cite this