TY - GEN
T1 - Fusion Pruning for Large Language Models
AU - Jiang, Shixin
AU - Liu, Ming
AU - Qin, Bing
N1 - Publisher Copyright:
©2024 IEEE.
PY - 2024
Y1 - 2024
N2 - Large language models have achieved great success in natural language processing tasks. It has recently become a new research hot spot. For example, in tasks such as mathematical reasoning and story writing, large models have emerged with extremely strong capabilities. However, their huge size and computing requirements have brought great challenges to actual deployment. In terms of reasoning speed, as the model size increases significantly, the model’s reasoning speed will drop a lot. Therefore, it is necessary to prune and accelerate large models. Existing structured and unstructured pruning methods have problems in compatibility and are not fully applicable to large models after pruning. Although these pruning methods are effective in theory, they usually show different applicability and effects when applied to complex models. For example, structured pruning methods may be more suitable for achieving model compression through sparse word embedding matrices or reducing the number of attention heads, while unstructured pruning methods focus more on pruning redundant parameter connections. However, these methods often lack sufficient compatibility and general applicability in practice. We mainly explore the research on fusion algorithms of pruning methods, including fusion acceleration solutions that combine structured pruning and unstructured pruning, as well as fusion acceleration solutions that combine pruning and other acceleration methods.
AB - Large language models have achieved great success in natural language processing tasks. It has recently become a new research hot spot. For example, in tasks such as mathematical reasoning and story writing, large models have emerged with extremely strong capabilities. However, their huge size and computing requirements have brought great challenges to actual deployment. In terms of reasoning speed, as the model size increases significantly, the model’s reasoning speed will drop a lot. Therefore, it is necessary to prune and accelerate large models. Existing structured and unstructured pruning methods have problems in compatibility and are not fully applicable to large models after pruning. Although these pruning methods are effective in theory, they usually show different applicability and effects when applied to complex models. For example, structured pruning methods may be more suitable for achieving model compression through sparse word embedding matrices or reducing the number of attention heads, while unstructured pruning methods focus more on pruning redundant parameter connections. However, these methods often lack sufficient compatibility and general applicability in practice. We mainly explore the research on fusion algorithms of pruning methods, including fusion acceleration solutions that combine structured pruning and unstructured pruning, as well as fusion acceleration solutions that combine pruning and other acceleration methods.
KW - Large language models
KW - model pruning
KW - model quantization
UR - https://www.scopus.com/pages/publications/85216389347
U2 - 10.1109/ISCSLP63861.2024.10800592
DO - 10.1109/ISCSLP63861.2024.10800592
M3 - 会议稿件
AN - SCOPUS:85216389347
T3 - 2024 14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024
SP - 349
EP - 352
BT - 2024 14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024
A2 - Qian, Yanmin
A2 - Jin, Qin
A2 - Ou, Zhijian
A2 - Ling, Zhenhua
A2 - Wu, Zhiyong
A2 - Li, Ya
A2 - Xie, Lei
A2 - Tao, Jianhua
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 14th International Symposium on Chinese Spoken Language Processing, ISCSLP 2024
Y2 - 7 November 2024 through 10 November 2024
ER -