Skip to main navigation Skip to search Skip to main content

Optimization of Power Allocation in Ultra-Dense Small Cellular Networks With Information-Gain-Guided Reinforcement Learning

  • Mingjian Fu
  • , Zhifei Lin
  • , Fenglin Ni
  • , Siyuan Li
  • , Xinyi Lin
  • , Zhenghong Wang*
  • , Guangyong Chen
  • , Yuanlong Yu
  • *Corresponding author for this work
  • Fuzhou University
  • Faculty of Computing, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

The exponential growth of data traffic in 6G drives the dense deployment of small cells. While spectrum reuse enhances capacity, it inevitably exacerbates inter-user and inter-cell interference. Consequently, power allocation is vital for maximizing the instantaneous downlink sum rate subject to transmit power constraints. However, traditional optimization methods, such as fractional programming, rely on global Channel State Information (CSI) and incur high computational complexity, rendering them impractical for dynamic, large scale networks. Although Deep Reinforcement Learning (DRL) offers a promising alternative, existing methods are limited by full CSI dependence, misalignment with instantaneous objectives, and inefficient exploration, hindering robustness in dynamic environments. To address these challenges, this study proposes an enhanced DRL-based power allocation framework termed Information-Gain Deep Deterministic Policy Gradient (IG-DDPG). First, a compact state representation is designed by retaining only the top-I strongest normalized interference channel gains relative to the serving link, while discarding weak and negligible interference sources. This selection rationale effectively preserves dominant interference information to maintain decision accuracy while significantly reducing the state space dimensionality to remove the dependency on global CSI. Second, a max-reward learning mechanism is incorporated into the Deep Deterministic Policy Gradient (DDPG) architecture to directly optimize instantaneous sum rate. Third, building upon the Intrinsic Curiosity Module (ICM), we reinterpret its prediction-error signal as an information-gain surrogate and integrate it as an intrinsic objective. A weighting parameter is introduced to balance the trade-off between the extrinsic instantaneous rate and the intrinsic information gain, guiding the agent toward dynamically informative state transitions. Extensive experiments demonstrate that IG-DDPG consistently outperforms the baseline DDPG algorithm. It improves the average sum rate by no less than 4.3%, increases the top 20% average sum rate by at least 2.4%, and reduces variance by up to 22.15%. These results confirm that IG-DDPG provides a scalable and robust solution for real-time interference management in 6G ultra-dense cellular networks.

Original languageEnglish
Pages (from-to)4419-4430
Number of pages12
JournalIEEE Transactions on Consumer Electronics
Volume72
Issue number2
DOIs
StatePublished - 1 May 2026
Externally publishedYes

Keywords

  • 6G
  • Power allocation
  • deep reinforcement learning
  • information-gain
  • ultra-dense cellular networks

Fingerprint

Dive into the research topics of 'Optimization of Power Allocation in Ultra-Dense Small Cellular Networks With Information-Gain-Guided Reinforcement Learning'. Together they form a unique fingerprint.

Cite this