TY - GEN
T1 - Age-Driven Joint Optimization for UAV Swarms via Multi-Agent Reinforcement Learning
AU - Wu, Haoxu
AU - Wu, Shaohua
AU - Meng, Siqi
AU - Tong, Yuze
AU - Zhang, Qinyu
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Real-time monitoring in remote areas presents critical challenges for time-sensitive missions such as emergency response and disaster surveillance, necessitating autonomous UAV swarms with stringent information timeliness requirements. For this end, we adopt a Leader-Follower UAV swarm architecture, in which Follower UAVs act as both sensing platforms and communication relays, thereby establishing a fully airborne network without reliance on ground infrastructure. To enhance information timeliness, we adopt an age-driven formulation quantified by the Age of Information (AoI), within which sampling decisions, buffer scheduling, and routing selection are jointly optimized under dynamic topology and partial observability. Accordingly, we model this problem as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP), and develop an enhanced multi-agent reinforcement learning (MARL) algorithm that employs specialized policy networks to decompose coupled decision-making into coordinated sub-tasks. Simulation results demonstrate the proposed approach significantly reduces average AoI compared to baseline algorithms while maintaining competitive performance in conventional network metrics including packet delivery ratio (PDR) and throughput.
AB - Real-time monitoring in remote areas presents critical challenges for time-sensitive missions such as emergency response and disaster surveillance, necessitating autonomous UAV swarms with stringent information timeliness requirements. For this end, we adopt a Leader-Follower UAV swarm architecture, in which Follower UAVs act as both sensing platforms and communication relays, thereby establishing a fully airborne network without reliance on ground infrastructure. To enhance information timeliness, we adopt an age-driven formulation quantified by the Age of Information (AoI), within which sampling decisions, buffer scheduling, and routing selection are jointly optimized under dynamic topology and partial observability. Accordingly, we model this problem as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP), and develop an enhanced multi-agent reinforcement learning (MARL) algorithm that employs specialized policy networks to decompose coupled decision-making into coordinated sub-tasks. Simulation results demonstrate the proposed approach significantly reduces average AoI compared to baseline algorithms while maintaining competitive performance in conventional network metrics including packet delivery ratio (PDR) and throughput.
KW - Age of information (AoI)
KW - multi-agent reinforcement learning (MARL)
KW - Real-time monitoring
KW - UAV swarms
UR - https://www.scopus.com/pages/publications/105042848672
U2 - 10.1109/WCNC65185.2026.11555621
DO - 10.1109/WCNC65185.2026.11555621
M3 - 会议稿件
AN - SCOPUS:105042848672
T3 - IEEE Wireless Communications and Networking Conference, WCNC
BT - 2026 IEEE Wireless Communications and Networking Conference, WCNC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2026 IEEE Wireless Communications and Networking Conference, WCNC 2026
Y2 - 13 April 2026 through 16 April 2026
ER -