Skip to main navigation Skip to search Skip to main content

Transforming Wikipedia Into Augmented Data for Query-Focused Summarization

  • Haichao Zhu
  • , Li Dong
  • , Furu Wei
  • , Bing Qin*
  • , Ting Liu
  • *Corresponding author for this work
  • Faculty of Computing, Harbin Institute of Technology
  • Microsoft USA

Research output: Contribution to journalArticlepeer-review

Abstract

The limited size of existing query-focused summarization datasets renders training data-driven summarization models challenging. Meanwhile, the manual construction of a query-focused summarization corpus is costly and time-consuming. In this paper, we use Wikipedia to automatically collect a large query-focused summarization dataset (named WikiRef) of more than 280,000 examples, which can serve as a means of data augmentation. We also develop a BERT-based query-focused summarization model (Q-BERT) to extract sentences from the documents as summaries. To better adapt a huge model containing millions of parameters to tiny benchmarks, we identify and fine-tune only a sparse subnetwork, which corresponds to a small fraction of the whole model parameters. Experimental results on three DUC benchmarks show that the model pre-trained on WikiRef has already achieved reasonable performance. After fine-tuning on the specific benchmark datasets, the model with data augmentation outperforms strong comparison systems. Moreover, both our proposed Q-BERT model and subnetwork fine-tuning further improve the model performance.

Original languageEnglish
Pages (from-to)2357-2367
Number of pages11
JournalIEEE/ACM Transactions on Audio Speech and Language Processing
Volume30
DOIs
StatePublished - 2022
Externally publishedYes

Keywords

  • Query-focused summarization
  • data augmentation
  • natural language processing
  • neural networks

Fingerprint

Dive into the research topics of 'Transforming Wikipedia Into Augmented Data for Query-Focused Summarization'. Together they form a unique fingerprint.

Cite this