TY - GEN
T1 - Battery Management for Warehouse Robots via Average-Reward Reinforcement Learning
AU - Mu, Yongjin
AU - Li, Yanjie
AU - Lin, Ke
AU - Deng, Ki
AU - Liu, Qi
N1 - Publisher Copyright:
© 2022 IEEE.
PY - 2022
Y1 - 2022
N2 - In automated warehouses, the battery management strategy of Automated Guided Vehicles (AGVs) can affect the throughput and operational efficiency of the warehouse. In this paper, we first model the battery management problem as a Markov Decision Process (MDP) and adopt the deep reinforcement learning (DRL) algorithm as the battery management strategy. However, discounted reward DRL algorithms ignore long-term benefits, which are not suitable for the strategy since orders arriving at the warehouse at every moment are important and should be treated. In order to solve the above problems, we then introduce the average reward DRL algorithm to focus more on long-term benefits. But the existing average reward DRL algorithms have the problems of low sample utilization and unstable training. Therefore, we present a practical algorithm called average reward TD3 (ARTD3) that learns faster and is more stable. Finally, we conduct extensive experiments to confirm that ARTD3 outperforms discounted reward DRL algorithm and rule-based methods.
AB - In automated warehouses, the battery management strategy of Automated Guided Vehicles (AGVs) can affect the throughput and operational efficiency of the warehouse. In this paper, we first model the battery management problem as a Markov Decision Process (MDP) and adopt the deep reinforcement learning (DRL) algorithm as the battery management strategy. However, discounted reward DRL algorithms ignore long-term benefits, which are not suitable for the strategy since orders arriving at the warehouse at every moment are important and should be treated. In order to solve the above problems, we then introduce the average reward DRL algorithm to focus more on long-term benefits. But the existing average reward DRL algorithms have the problems of low sample utilization and unstable training. Therefore, we present a practical algorithm called average reward TD3 (ARTD3) that learns faster and is more stable. Finally, we conduct extensive experiments to confirm that ARTD3 outperforms discounted reward DRL algorithm and rule-based methods.
KW - Automated Warehouses
KW - Average-Reward Reinforcement Learning
KW - Battery Management
UR - https://www.scopus.com/pages/publications/85147330666
U2 - 10.1109/ROBIO55434.2022.10011784
DO - 10.1109/ROBIO55434.2022.10011784
M3 - 会议稿件
AN - SCOPUS:85147330666
T3 - 2022 IEEE International Conference on Robotics and Biomimetics, ROBIO 2022
SP - 253
EP - 258
BT - 2022 IEEE International Conference on Robotics and Biomimetics, ROBIO 2022
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2022 IEEE International Conference on Robotics and Biomimetics, ROBIO 2022
Y2 - 5 December 2022 through 9 December 2022
ER -