Abstract
Deterministic policy gradient methods with neural network actors, such as Twin Delayed Deep Deterministic Policy Gradient (TD3), offer strong performance but remain difficult to interpret. Although existing fuzzy reinforcement learning methods can improve interpretability to some extent, embedding a fully interpretable fuzzy actor into the TD3 pipeline with rule-level prior knowledge remains challenging. To address these challenges, we propose Fuzzy–TD3, a deterministic actor–critic algorithm that replaces the neural actor with a Takagi–Sugeno (T–S) fuzzy policy and integrates prior knowledge at the rule level. In this framework, the consequents of fuzzy rules associated with locally linearized operating regions are fixed to an optimal feedback law, while the remaining rule consequents are learned directly from data. Antecedents are reparameterized to maintain valid fuzzy partitions and are trained via gradient descent, ensuring interpretability. Simulation results on two inverted-pendulum benchmarks show that Fuzzy–TD3 improves sample efficiency and steady-state regulation while maintaining competitive transient and input-usage performance compared with neural actor–critic baselines. This work provides an interpretable and practical reinforcement learning framework that unites fuzzy theory with classical control, offering a robust solution for data-driven control in industrial applications.
| Original language | English |
|---|---|
| Journal | ISA Transactions |
| DOIs | |
| State | Accepted/In press - 2026 |
| Externally published | Yes |
Keywords
- Deterministic policy gradient
- Interpretability
- Optimal control
- Prior knowledge
- Sample efficiency
- Takagi–Sugeno fuzzy model
Fingerprint
Dive into the research topics of 'Fuzzy–TD3: A prior-knowledge guided deterministic policy gradient for interpretable control'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver