Abstract
The standard paradigm of reinforcement learning (RL) is the Markov Decision Process (MDP) in which an agent learns to maximize the cumulative discounted rewards. The reward function in MDP is generally defined as the sum of multiple reward components, each designed to encapsulate a specific aspect of the expected policy. The discount factor γ ∈ [0, 1) decreases the future reward in the present value, which determines the effective time horizon for the agent. In the conventional MDP, all reward components are subject to the same discount factor regardless of their specific meanings. Although this convenient configuration simplifies the problem in the algorithm deployment, it sacrifices precision in defining the optimization problem and results in a temporal mismatch of rewards with diverse physical meanings. This paper proposes multi-discounting MDP (MDMDP), a novel model based on reward decomposition to solve the above problems. MDMDP allows practitioners to set separate discount factors for different reward components. This capability provides great flexibility in combining reward components at different timescales. Furthermore, this paper proposes an RL algorithm, multi-discounting Q-learning, to solve finite MDMDP. Moreover, we extend it to the deep RL version, including multi-discounting DQN for discrete action space tasks and multi-discounting actor-critic for continuous action space tasks. Experimental results demonstrate that the proposed methods improve flexibility and precision in modeling complex tasks, enhancing the alignment of the agent's policy with desired objectives.
| Original language | English |
|---|---|
| Pages (from-to) | 94-104 |
| Number of pages | 11 |
| Journal | IEEE Transactions on Emerging Topics in Computing |
| Volume | 14 |
| Issue number | 1 |
| DOIs | |
| State | Published - 2026 |
| Externally published | Yes |
Keywords
- MDP
- discount factor
- reinforcement learning
- reward decomposition
Fingerprint
Dive into the research topics of 'Multi-Discounting Reinforcement Learning Based on Reward Decomposition'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver