Skip to main navigation Skip to search Skip to main content

Distilling Structured Rationale from Large Language Models to Small Language Models for Abstractive Summarization

  • Linyong Wang
  • , Lianwei Wu*
  • , Shaoqi Song
  • , Yaxiong Wang
  • , Cuiyun Gao
  • , Kang Wang
  • *Corresponding author for this work
  • Northwestern Polytechnical University Xian
  • Hefei University of Technology
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalConference articlepeer-review

Abstract

Large Language Models (LLMs) have permeated various Natural Language Processing (NLP) tasks. For the summarization tasks, LLMs can generate well-structured rationales, which consist of Essential Aspects (EA), Associated Sentences (AS) and Triple Entity Relations (TER). These rationales guide smaller models (=1B) to produce better summaries. However, their high deployment costs (=70B), such as substantial storage space and high computing requirements, limit their utilization in resource-constrained environments. Furthermore, effectively distilling these structured rationales from LLMs into Small Language Models (SLMs) models remains a challenge. To address this, we propose the LLM-based Structured Rationale-guided Multi-view Weak-gated Fusion framework (LSR-MWF). The framework initially employs LLMs to dig structural rationales from a document, considering multiple viewpoints such as EA, AS, and TER. Then, it develop a multi-step summary generation evaluation strategy to select high-quality structured rationales. Subsequently, it aligns with these rationales using additional modules organized in a hierarchical structure. Finally, the framework integrates the features output by these modules with original abstractive model through a weak-gated mechanism. Experimental results on two publicly available CNN/DailyMail and XSum datasets show that our method improves the performance of the abstractive model, outperforming baselines by 11.2% and 5.8%, respectively. In addition, our method improves the interpretability of summary generation from the viewpoints of EA, AS and TER.

Original languageEnglish
Pages (from-to)25389-25397
Number of pages9
JournalProceedings of the AAAI Conference on Artificial Intelligence
Volume39
Issue number24
DOIs
StatePublished - 11 Apr 2025
Externally publishedYes
Event39th Annual AAAI Conference on Artificial Intelligence, AAAI 2025 - Philadelphia, United States
Duration: 25 Feb 20254 Mar 2025

Fingerprint

Dive into the research topics of 'Distilling Structured Rationale from Large Language Models to Small Language Models for Abstractive Summarization'. Together they form a unique fingerprint.

Cite this