Skip to main navigation Skip to search Skip to main content

Fusion Pruning for Large Language Models

  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Large language models have achieved great success in natural language processing tasks. It has recently become a new research hot spot. For example, in tasks such as mathematical reasoning and story writing, large models have emerged with extremely strong capabilities. However, their huge size and computing requirements have brought great challenges to actual deployment. In terms of reasoning speed, as the model size increases significantly, the model’s reasoning speed will drop a lot. Therefore, it is necessary to prune and accelerate large models. Existing structured and unstructured pruning methods have problems in compatibility and are not fully applicable to large models after pruning. Although these pruning methods are effective in theory, they usually show different applicability and effects when applied to complex models. For example, structured pruning methods may be more suitable for achieving model compression through sparse word embedding matrices or reducing the number of attention heads, while unstructured pruning methods focus more on pruning redundant parameter connections. However, these methods often lack sufficient compatibility and general applicability in practice. We mainly explore the research on fusion algorithms of pruning methods, including fusion acceleration solutions that combine structured pruning and unstructured pruning, as well as fusion acceleration solutions that combine pruning and other acceleration methods.

Original languageEnglish
Title of host publication2024 14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024
EditorsYanmin Qian, Qin Jin, Zhijian Ou, Zhenhua Ling, Zhiyong Wu, Ya Li, Lei Xie, Jianhua Tao
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages349-352
Number of pages4
ISBN (Electronic)9798331516826
DOIs
StatePublished - 2024
Event14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024 - Beijing, China
Duration: 7 Nov 202410 Nov 2024

Publication series

Name2024 14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024

Conference

Conference14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024
Country/TerritoryChina
CityBeijing
Period7/11/2410/11/24

Keywords

  • Large language models
  • model pruning
  • model quantization

Fingerprint

Dive into the research topics of 'Fusion Pruning for Large Language Models'. Together they form a unique fingerprint.

Cite this