Skip to main navigation Skip to search Skip to main content

A deep reinforcement learning state representation method that integrates compression and iterative external memory networks

  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Deep reinforcement learning (DRL) faces short-term sequence modeling difficulties and long-term credit assignment problems (such as reward sparsity and memory challenges caused by delays) when extracting spatiotemporal features from high-dimensional pixel inputs. We propose a deep reinforcement learning framework based on compressed iterative external memory networks. The core contribution of this work is a compression-and-iterative external memory framework for visual reinforcement learning, in which compact spatiotemporal encoding, memory-efficient storage, and multi-hop memory reasoning are jointly designed for long-horizon decision-making. Specifically, the method first compresses stacked image observations into low-dimensional spatiotemporal latent queries using 3D CNNs and RoPE-enhanced self-attention, and then performs iterative multi-hop reasoning over a fixed-size external memory module with two-stage memory compression. Experiments were conducted in the Atari (discrete action space) and MuJoCo (continuous action space) benchmark environments, and the results show that the proposed method significantly improves sample efficiency and agent performance, especially in tasks requiring spatiotemporal perception and long-term planning.

Original languageEnglish
Article number134169
JournalNeurocomputing
Volume697
DOIs
StatePublished - 7 Oct 2026
Externally publishedYes

Keywords

  • Deep reinforcement learning
  • Memory iteration
  • Memory vector compression
  • Rotary position embedding
  • Spatiotemporal reasoning

Fingerprint

Dive into the research topics of 'A deep reinforcement learning state representation method that integrates compression and iterative external memory networks'. Together they form a unique fingerprint.

Cite this