TY - GEN
T1 - COPR
T2 - 63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
AU - Zhang, Han
AU - Gui, Lin
AU - Lei, Yu
AU - Zhai, Yuanzhao
AU - Zhang, Yehong
AU - Zhang, Zhuo
AU - He, Yulan
AU - Wang, Hui
AU - Yu, Yue
AU - Wong, Kam Fai
AU - Liang, Bin
AU - Xu, Ruifeng
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics.
PY - 2025
Y1 - 2025
N2 - The growing integration of Large Language Models (LLMs) into real-world applications underscores the critical need for continual alignment with evolving human preferences. Reinforcement Learning from Human Feedback (RLHF) has shown success in improving the alignment of LLMs, but its rigid, multi-stage process presents significant limitations for continual learning (CL) scenarios, where models need to adapt incrementally without catastrophic forgetting. Existing methods, such as Direct Preference Optimization (DPO), offer potential for offline preference learning but exhibit challenges like increased preference gap amplification and reduced model diversity, which can lead to preference collapse. In practical settings, LLMs continuously interact with diverse user feedback across tasks and domains. The inability of current approaches to efficiently incorporate incremental human preferences without retraining or significant computational overhead limits their scalability and adaptability. Addressing these gaps, our study introduces a novel framework, Continual Optimal Policy Regularization (COPR), that ensures robust and flexible continual alignment while preserving historical knowledge and optimizing performance in new preference tasks.
AB - The growing integration of Large Language Models (LLMs) into real-world applications underscores the critical need for continual alignment with evolving human preferences. Reinforcement Learning from Human Feedback (RLHF) has shown success in improving the alignment of LLMs, but its rigid, multi-stage process presents significant limitations for continual learning (CL) scenarios, where models need to adapt incrementally without catastrophic forgetting. Existing methods, such as Direct Preference Optimization (DPO), offer potential for offline preference learning but exhibit challenges like increased preference gap amplification and reduced model diversity, which can lead to preference collapse. In practical settings, LLMs continuously interact with diverse user feedback across tasks and domains. The inability of current approaches to efficiently incorporate incremental human preferences without retraining or significant computational overhead limits their scalability and adaptability. Addressing these gaps, our study introduces a novel framework, Continual Optimal Policy Regularization (COPR), that ensures robust and flexible continual alignment while preserving historical knowledge and optimizing performance in new preference tasks.
UR - https://www.scopus.com/pages/publications/105028587500
U2 - 10.18653/v1/2025.findings-acl.281
DO - 10.18653/v1/2025.findings-acl.281
M3 - 会议稿件
AN - SCOPUS:105028587500
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 5377
EP - 5398
BT - Findings of the Association for Computational Linguistics
A2 - Che, Wanxiang
A2 - Nabende, Joyce
A2 - Shutova, Ekaterina
A2 - Pilehvar, Mohammad Taher
PB - Association for Computational Linguistics (ACL)
Y2 - 27 July 2025 through 1 August 2025
ER -