Abstract
The exponential growth of data traffic in 6G drives the dense deployment of small cells. While spectrum reuse enhances capacity, it inevitably exacerbates inter-user and inter-cell interference. Consequently, power allocation is vital for maximizing the instantaneous downlink sum rate subject to transmit power constraints. However, traditional optimization methods, such as fractional programming, rely on global Channel State Information (CSI) and incur high computational complexity, rendering them impractical for dynamic, large scale networks. Although Deep Reinforcement Learning (DRL) offers a promising alternative, existing methods are limited by full CSI dependence, misalignment with instantaneous objectives, and inefficient exploration, hindering robustness in dynamic environments. To address these challenges, this study proposes an enhanced DRL-based power allocation framework termed Information-Gain Deep Deterministic Policy Gradient (IG-DDPG). First, a compact state representation is designed by retaining only the top-I strongest normalized interference channel gains relative to the serving link, while discarding weak and negligible interference sources. This selection rationale effectively preserves dominant interference information to maintain decision accuracy while significantly reducing the state space dimensionality to remove the dependency on global CSI. Second, a max-reward learning mechanism is incorporated into the Deep Deterministic Policy Gradient (DDPG) architecture to directly optimize instantaneous sum rate. Third, building upon the Intrinsic Curiosity Module (ICM), we reinterpret its prediction-error signal as an information-gain surrogate and integrate it as an intrinsic objective. A weighting parameter is introduced to balance the trade-off between the extrinsic instantaneous rate and the intrinsic information gain, guiding the agent toward dynamically informative state transitions. Extensive experiments demonstrate that IG-DDPG consistently outperforms the baseline DDPG algorithm. It improves the average sum rate by no less than 4.3%, increases the top 20% average sum rate by at least 2.4%, and reduces variance by up to 22.15%. These results confirm that IG-DDPG provides a scalable and robust solution for real-time interference management in 6G ultra-dense cellular networks.
| Original language | English |
|---|---|
| Pages (from-to) | 4419-4430 |
| Number of pages | 12 |
| Journal | IEEE Transactions on Consumer Electronics |
| Volume | 72 |
| Issue number | 2 |
| DOIs | |
| State | Published - 1 May 2026 |
| Externally published | Yes |
Keywords
- 6G
- Power allocation
- deep reinforcement learning
- information-gain
- ultra-dense cellular networks
Fingerprint
Dive into the research topics of 'Optimization of Power Allocation in Ultra-Dense Small Cellular Networks With Information-Gain-Guided Reinforcement Learning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver