Skip to main navigation Skip to search Skip to main content

COPR: Continual Human Preference Learning via Optimal Policy Regularization

  • Han Zhang
  • , Lin Gui
  • , Yu Lei
  • , Yuanzhao Zhai
  • , Yehong Zhang
  • , Zhuo Zhang
  • , Yulan He
  • , Hui Wang
  • , Yue Yu
  • , Kam Fai Wong
  • , Bin Liang*
  • , Ruifeng Xu*
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • Peng Cheng Laboratory
  • King's College London
  • National University of Defense Technology
  • Chinese University of Hong Kong
  • Guangdong Provincial Key Laboratory of Novel Security Intelligence Technologies

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The growing integration of Large Language Models (LLMs) into real-world applications underscores the critical need for continual alignment with evolving human preferences. Reinforcement Learning from Human Feedback (RLHF) has shown success in improving the alignment of LLMs, but its rigid, multi-stage process presents significant limitations for continual learning (CL) scenarios, where models need to adapt incrementally without catastrophic forgetting. Existing methods, such as Direct Preference Optimization (DPO), offer potential for offline preference learning but exhibit challenges like increased preference gap amplification and reduced model diversity, which can lead to preference collapse. In practical settings, LLMs continuously interact with diverse user feedback across tasks and domains. The inability of current approaches to efficiently incorporate incremental human preferences without retraining or significant computational overhead limits their scalability and adaptability. Addressing these gaps, our study introduces a novel framework, Continual Optimal Policy Regularization (COPR), that ensures robust and flexible continual alignment while preserving historical knowledge and optimizing performance in new preference tasks.

Original languageEnglish
Title of host publicationFindings of the Association for Computational Linguistics
Subtitle of host publicationACL 2025
EditorsWanxiang Che, Joyce Nabende, Ekaterina Shutova, Mohammad Taher Pilehvar
PublisherAssociation for Computational Linguistics (ACL)
Pages5377-5398
Number of pages22
ISBN (Electronic)9798891762565
DOIs
StatePublished - 2025
Event63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025 - Vienna, Austria
Duration: 27 Jul 20251 Aug 2025

Publication series

NameProceedings of the Annual Meeting of the Association for Computational Linguistics
ISSN (Print)0736-587X

Conference

Conference63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
Country/TerritoryAustria
CityVienna
Period27/07/251/08/25

Fingerprint

Dive into the research topics of 'COPR: Continual Human Preference Learning via Optimal Policy Regularization'. Together they form a unique fingerprint.

Cite this