Abstract
Scientific papers follow structured facets (e.g., Introduction, Methods), and modern research dissemination increasingly incorporates multimodal formats like presentation videos and audio. This shift necessitates summarization systems that can process both structured and multimodal information. This work proposes Localize-then-Summarize (LENS), a two-stage scientific summarizer that first localizes relevant presentation segments that align with paper facets; followed by summarizing them via memory-augmented reasoning that models dependencies across modalities and facets. On a new MFS-SciSum dataset with 2.7k aligned paper–presentation pairs, the LENS localizer and summarizer achieve Recall@0.5/0.7 scores of 40.83/23.06, and ROUGE-1/2/L scores of 44.71/15.26/21.64, outperforming strong baselines like CLIP and Transformer by 10–15 points in Recall and 1–5 points in ROUGE (resp.). Additionally, the LENS summarizer reduces GPU usage by 71%, notably improving generation efficiency. Code and data are available at: https://github.com/allent4n/LENS.
| Original language | English |
|---|---|
| Article number | 104987 |
| Journal | Information Processing and Management |
| Volume | 64 |
| Issue number | 1 |
| DOIs | |
| State | Published - Jan 2027 |
| Externally published | Yes |
Keywords
- Cross-Modal Facet Localizer
- Faceted summarization of scientific documents
- Memory-enhanced summarizer
- Multimodal summarization
Fingerprint
Dive into the research topics of 'Localize-then-summarize: Enhancing scientific multimodal summarization with facet-aware cross-modal memory'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver