TY - GEN
T1 - Disentangled Task Representation Learning for Offline Meta Reinforcement Learning
AU - Cong, Shan
AU - Yu, Chao
AU - Wang, Yaowei
AU - Jiang, Dongmei
AU - Lan, Xiangyuan
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - In this paper, we aim to address the generalization problem in Offline Meta-Reinforcement Learning (OMRL) when both task objectives and environmental parameters vary simultaneously. We propose DIsentangled TAsk Representation learning (DITAR) for OMRL, which leverages the Conditional Variational Auto-Encoder framework to disentangle task representations into distinct components for task objectives and environmental parameters, thus enhancing policy generalization across diverse tasks. We further impose orthogonality constraints to prevent overlap between the representation spaces, ensuring that each space independently captures its corresponding task component. Additionally, mutual information optimization is applied to remove redundant state-action information, focusing the representation space on task-relevant features. Experiments on the multi-task MuJoCo benchmark demonstrate that DITAR significantly outperforms previous methods, particularly in complex scenarios involving simultaneous changes in both task objectives and environmental parameters.
AB - In this paper, we aim to address the generalization problem in Offline Meta-Reinforcement Learning (OMRL) when both task objectives and environmental parameters vary simultaneously. We propose DIsentangled TAsk Representation learning (DITAR) for OMRL, which leverages the Conditional Variational Auto-Encoder framework to disentangle task representations into distinct components for task objectives and environmental parameters, thus enhancing policy generalization across diverse tasks. We further impose orthogonality constraints to prevent overlap between the representation spaces, ensuring that each space independently captures its corresponding task component. Additionally, mutual information optimization is applied to remove redundant state-action information, focusing the representation space on task-relevant features. Experiments on the multi-task MuJoCo benchmark demonstrate that DITAR significantly outperforms previous methods, particularly in complex scenarios involving simultaneous changes in both task objectives and environmental parameters.
KW - conditional variational auto-encoder
KW - meta-policy optimization
KW - mutual information mechanism
KW - task representation learning
UR - https://www.scopus.com/pages/publications/85215629058
U2 - 10.1109/ICA63002.2024.00022
DO - 10.1109/ICA63002.2024.00022
M3 - 会议稿件
AN - SCOPUS:85215629058
T3 - Proceedings - 2024 IEEE International Conference on Agents, ICA 2024
SP - 64
EP - 69
BT - Proceedings - 2024 IEEE International Conference on Agents, ICA 2024
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2024 IEEE International Conference on Agents, ICA 2024
Y2 - 4 December 2024 through 6 December 2024
ER -