Abstract
A reinforcement learning algorithm based on linear average is proposed, which is used to solve non-convergent problems of reinforcement learning function approximation in continuous state space. According to contraction theory, this algorithm is based on gradient descent method, which adopts linear average as performance evaluation of value function. So the iterative process of value function becomes a convergent process to a fixed value. A standard reinforcement learning problem, Mountain Car Problem, is used to verify the performance of the algorithm. Results show the effectiveness, feasibility and quick convergence of the algorithm.
| Original language | English |
|---|---|
| Pages (from-to) | 1407-1411 |
| Number of pages | 5 |
| Journal | Jilin Daxue Xuebao (Gongxueban)/Journal of Jilin University (Engineering and Technology Edition) |
| Volume | 38 |
| Issue number | 6 |
| State | Published - Nov 2008 |
| Externally published | Yes |
Keywords
- Automatic control technology
- Function approximation
- Gradient descent method
- Linear averages
- Reinforcement learning
Fingerprint
Dive into the research topics of 'Reinforcement learning function approximation algorithm based on linear average'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver