Abstract
Deep reinforcement learning algorithms often require large amounts of training data, particularly in robotic control tasks. To address this limitation, we propose a sample-efficient backtrack temporal difference learning method that enhances target state-action (Q) value estimation. The proposed method dynamically prioritizes transitions based on their proximity to terminal states using backtrack sampling weights. This prioritization mechanism yields more accurate target Q-values, thereby improving the overall Q-value estimation precision. Furthermore, our analysis uncovers a novel link between curriculum learning and Bellman equation optimization. The proposed method is versatile, applicable to both discrete and continuous action spaces, and readily integrable with off-policy actor-critic algorithms. Extensive experiments show that the proposed method considerably reduces Q-value approximation errors and outperforms baselines across diverse benchmarks, achieving a 28 % performance improvement in four discrete action-space tasks and a 78 % gain in four continuous control tasks.
| Original language | English |
|---|---|
| Article number | 114613 |
| Journal | Knowledge-Based Systems |
| Volume | 330 |
| DOIs | |
| State | Published - 25 Nov 2025 |
| Externally published | Yes |
Keywords
- Backtrack temporal difference
- Reinforcement learning
- Robot control
- Sample efficiency
Fingerprint
Dive into the research topics of 'Sample-efficient backtrack temporal difference deep reinforcement learning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver