Skip to main navigation Skip to search Skip to main content

Localize-then-summarize: Enhancing scientific multimodal summarization with facet-aware cross-modal memory

  • Zusheng Tan
  • , Jing Yu Ji
  • , Wenhui Yu
  • , Ngai Fung Ng
  • , Fan Yang
  • , Jeff Tang
  • , Ken Fong
  • , Jing Li
  • , Sam Kwong
  • , Billy Chiu*
  • *Corresponding author for this work
  • Lingnan University
  • Shenzhen University
  • Hong Kong Polytechnic University
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalArticlepeer-review

Abstract

Scientific papers follow structured facets (e.g., Introduction, Methods), and modern research dissemination increasingly incorporates multimodal formats like presentation videos and audio. This shift necessitates summarization systems that can process both structured and multimodal information. This work proposes Localize-then-Summarize (LENS), a two-stage scientific summarizer that first localizes relevant presentation segments that align with paper facets; followed by summarizing them via memory-augmented reasoning that models dependencies across modalities and facets. On a new MFS-SciSum dataset with 2.7k aligned paper–presentation pairs, the LENS localizer and summarizer achieve Recall@0.5/0.7 scores of 40.83/23.06, and ROUGE-1/2/L scores of 44.71/15.26/21.64, outperforming strong baselines like CLIP and Transformer by 10–15 points in Recall and 1–5 points in ROUGE (resp.). Additionally, the LENS summarizer reduces GPU usage by 71%, notably improving generation efficiency. Code and data are available at: https://github.com/allent4n/LENS.

Original languageEnglish
Article number104987
JournalInformation Processing and Management
Volume64
Issue number1
DOIs
StatePublished - Jan 2027
Externally publishedYes

Keywords

  • Cross-Modal Facet Localizer
  • Faceted summarization of scientific documents
  • Memory-enhanced summarizer
  • Multimodal summarization

Fingerprint

Dive into the research topics of 'Localize-then-summarize: Enhancing scientific multimodal summarization with facet-aware cross-modal memory'. Together they form a unique fingerprint.

Cite this