TY - GEN
T1 - Group Relative Policy Optimization for Robust Blind Interference Alignment with Fluid Antennas
AU - Peng, Jianqiu
AU - Zhang, Tong
AU - Wang, Shuai
AU - Shao, Mingjie
AU - Xu, Hao
AU - Wang, Rui
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Fluid antenna system (FAS) leverages dynamic re-configurability to unlock spatial degrees of freedom and reshape wireless channels. Blind interference alignment (BIA) aligns interference through antenna switching. This paper proposes, for the first time, a robust fluid antenna-driven BIA framework for a K-user MISO downlink under imperfect channel state information (CSI). We formulate a robust sum-rate maximization problem through optimizing fluid antenna positions (switching positions). To solve this challenging non-convex problem, we employ group relative policy optimization (GRPO), a novel deep reinforcement learning algorithm that eliminates the critic network. This robust design reduces model size and floating point operations (FLOPs) by nearly half compared to proximal policy optimization (PPO) while significantly enhancing performance through group-based exploration that escapes bad local optima. Simulation results demonstrate that GRPO outperforms PPO by 4.17%, and a 100K-step pre-trained PPO by 30.29%. Due to error distribution learning, GRPO exceeds heuristic MaximumGain and RandomGain by 200.78% and 465.38%, respectively.
AB - Fluid antenna system (FAS) leverages dynamic re-configurability to unlock spatial degrees of freedom and reshape wireless channels. Blind interference alignment (BIA) aligns interference through antenna switching. This paper proposes, for the first time, a robust fluid antenna-driven BIA framework for a K-user MISO downlink under imperfect channel state information (CSI). We formulate a robust sum-rate maximization problem through optimizing fluid antenna positions (switching positions). To solve this challenging non-convex problem, we employ group relative policy optimization (GRPO), a novel deep reinforcement learning algorithm that eliminates the critic network. This robust design reduces model size and floating point operations (FLOPs) by nearly half compared to proximal policy optimization (PPO) while significantly enhancing performance through group-based exploration that escapes bad local optima. Simulation results demonstrate that GRPO outperforms PPO by 4.17%, and a 100K-step pre-trained PPO by 30.29%. Due to error distribution learning, GRPO exceeds heuristic MaximumGain and RandomGain by 200.78% and 465.38%, respectively.
KW - Blind interference alignment
KW - fluid antenna system
KW - group relative policy optimization
KW - sum-rate
UR - https://www.scopus.com/pages/publications/105045359818
U2 - 10.1109/ICC59461.2026.11587674
DO - 10.1109/ICC59461.2026.11587674
M3 - 会议稿件
AN - SCOPUS:105045359818
T3 - IEEE International Conference on Communications
BT - ICC 2026 - IEEE International Conference on Communications, Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2026 IEEE International Conference on Communications, ICC 2026
Y2 - 24 May 2026 through 28 May 2026
ER -