Skip to main navigation Skip to search Skip to main content

Mitigating Constraint Conflict in Offline RL: An Adaptive Weighted Constraint Approach

  • Pengyu Chen*
  • , Shirong Liu
  • , Minye Huang
  • , Haozhuo Zheng
  • , Haoyu Liu
  • , Wenyu Yuan
  • , Yang Liu
  • *Corresponding author for this work
  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Offline Reinforcement Learning allows agents to learn policies from pre-collected datasets by imposing conservative constraints to address the out-of-distribution problem. However, existing methods face a critical challenge of constraint conflict with datasets generated by multiple behavior policies, and these policies suggest conflicting actions that lead the agent towards suboptimal performance. Geometric distance-based and advantage-weighted methods can be employed to address this problem. These methods exhibit several limitations, including sensitivity to low-quality data, high computational cost, and over conservatism. To overcome these limitations, we propose Adaptive Weighted Constraint (AWC), which mitigates constraint conflicts by training a constraint network via adaptive weighted behavior cloning. AWC dynamically assigns importance weights to dataset actions based on their consistency with the current policy, ensuring that the constraint is informed by the behavior and its distance to the policy. Inspired by the robustness of central tendency estimators in statistics, we apply the weighted geometric median of the actions as a stable target for the policy constraint. Experiments on D4RL benchmarks demonstrate AWC outperforms prior methods on a majority of tasks.

Original languageEnglish
Title of host publicationAAMAS 2026 - Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems
PublisherAssociation for Computing Machinery, Inc
Pages3621-3623
Number of pages3
ISBN (Electronic)9798400723179
DOIs
StatePublished - 24 May 2026
Event25th International Conference on Autonomous Agents and Multiagent Systems, AAMAS 2026 - Paphos, Cyprus
Duration: 25 May 202629 May 2026

Publication series

NameAAMAS 2026 - Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems

Conference

Conference25th International Conference on Autonomous Agents and Multiagent Systems, AAMAS 2026
Country/TerritoryCyprus
CityPaphos
Period25/05/2629/05/26

Keywords

  • Constraint Conflict
  • Geometric Median
  • Offline Reinforcement Learning

Fingerprint

Dive into the research topics of 'Mitigating Constraint Conflict in Offline RL: An Adaptive Weighted Constraint Approach'. Together they form a unique fingerprint.

Cite this