TY - GEN
T1 - TCN-SAC
T2 - 2025 China Automation Congress, CAC 2025
AU - Zhou, Yi
AU - Yang, Jiaming
AU - Li, Zijing
AU - Qiu, Jianbin
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - In the past few years, deep reinforcement learning (DRL) has garnered increasing attention and found widespread use in fields such as robot control, gaming, and autonomous driving. Among model-free reinforcement learning algorithms, the Soft Actor-Critic (SAC) algorithm stands out due to its exceptional exploration ability and high data efficiency, making it a popular choice in both research and practical applications. However, the traditional SAC algorithm lacks the ability to model temporal dependencies in historical state data. Moreover, the current improvements using Long Short-Term Memory (LSTM) as a solution to this limitation still face challenges, including low training efficiency and instability training. In this paper, the TCN-SAC algorithm is proposed, which employs Temporal Convolutional Networks (TCNs) as the backbone for the Actor network. This proposed algorithm models the temporal dependencies in historical states while offering advantages such as high parallelism, flexible receptive fields, and stable training. Additionally, comparative experiments conducted in the Pendulum-v1 environment demonstrate that the proposed TCN-SAC outperforms the baseline, highlighting the superiority of this algorithm.
AB - In the past few years, deep reinforcement learning (DRL) has garnered increasing attention and found widespread use in fields such as robot control, gaming, and autonomous driving. Among model-free reinforcement learning algorithms, the Soft Actor-Critic (SAC) algorithm stands out due to its exceptional exploration ability and high data efficiency, making it a popular choice in both research and practical applications. However, the traditional SAC algorithm lacks the ability to model temporal dependencies in historical state data. Moreover, the current improvements using Long Short-Term Memory (LSTM) as a solution to this limitation still face challenges, including low training efficiency and instability training. In this paper, the TCN-SAC algorithm is proposed, which employs Temporal Convolutional Networks (TCNs) as the backbone for the Actor network. This proposed algorithm models the temporal dependencies in historical states while offering advantages such as high parallelism, flexible receptive fields, and stable training. Additionally, comparative experiments conducted in the Pendulum-v1 environment demonstrate that the proposed TCN-SAC outperforms the baseline, highlighting the superiority of this algorithm.
KW - deep reinforcement learning
KW - soft Actor-Critic
KW - temporal convolutional networks
KW - temporal modeling
UR - https://www.scopus.com/pages/publications/105041001118
U2 - 10.1109/CAC67268.2025.11487489
DO - 10.1109/CAC67268.2025.11487489
M3 - 会议稿件
AN - SCOPUS:105041001118
T3 - Proceedings - 2025 China Automation Congress, CAC 2025
SP - 3600
EP - 3605
BT - Proceedings - 2025 China Automation Congress, CAC 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 26 September 2025 through 28 September 2025
ER -