Skip to main navigation Skip to search Skip to main content

Contextual Policy Transfer in Meta-Reinforcement Learning via Active Learning

  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In meta-reinforcement learning (meta-RL), agents that consider the context when transferring source policies have been shown to outperform context-free approaches. However, existing approaches require large amounts of on-policy experience to adapt to novel tasks, limiting their practicality and sample efficiency. In this paper, we jointly perform off-policy meta-RL and active learning to generate the latent context of the novel task by reusing valuable experiences from source tasks. To calculate the importance weight of source experience for adaptation, we employ maximum mean discrepancy (MMD) as the criterion to minimize the experience distribution distance between the target task and the adapted source tasks in a reproducing kernel Hilbert space (RKHS). Integrating source experiences based on active queries with a small amount of on-policy target experience, we demonstrate that the experience sampling benefits the fine-tuning of the contextual policy. Then, we incorporate it into a standard meta-RL framework and verify its effectiveness on four continuous control environments, simulated via the MuJoCo simulator.

Original languageEnglish
Title of host publicationWeb Information Systems and Applications - 19th International Conference, WISA 2022, Proceedings
EditorsXiang Zhao, Shiyu Yang, Xin Wang, Jianxin Li
PublisherSpringer Science and Business Media Deutschland GmbH
Pages354-365
Number of pages12
ISBN (Print)9783031203084
DOIs
StatePublished - 2022
Event19th International Conference on Web Information Systems and Applications, WISA 2022 - Dalian, China
Duration: 16 Sep 202218 Sep 2022

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume13579 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference19th International Conference on Web Information Systems and Applications, WISA 2022
Country/TerritoryChina
CityDalian
Period16/09/2218/09/22

Keywords

  • Active learning
  • Meta reinforcement learning
  • MuJoCo
  • Uncertain sampling

Fingerprint

Dive into the research topics of 'Contextual Policy Transfer in Meta-Reinforcement Learning via Active Learning'. Together they form a unique fingerprint.

Cite this