Abstract
Implicit Neural Representations (INRs) have emerged as a promising paradigm for video coding by enabling continuous signal modeling through coordinate-based neural networks. However, existing video INR methods often model temporal variation either in the pixel domain or with a single dynamic code, which limits their ability to jointly preserve rapidly varying high-frequency details and long-range semantic motion patterns. To address this issue, we propose the Latent-disentangled Neural Representation for Videos (LNeRV), a framework that models video dynamics in latent space using dual dynamic feature grids. Specifically, a denser high-resolution grid is used to capture short-term detail dynamics, while a coarser low-rate grid is used to capture long-range semantic motion patterns. On top of these complementary grids, we introduce a Clustering-based Context Aggregation (CCA) module to combine short-term motion details with long-range temporal context, and a Decoupled Gated Attention (DGA) module as a lightweight gated fusion mechanism to enlarge the receptive field during feature interaction. Extensive experiments demonstrate that LNeRV achieves state-of-the-art video regression performance on the UVG dataset under both PSNR and MS-SSIM metrics. Compared with vanilla NeRV, LNeRV achieves an average PSNR gain of 2.52 dB on UVG, while also showing strong structural fidelity in the compressed rate-distortion setting.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Circuits and Systems for Video Technology |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- Image Processing
- Implicit neural representation (INR)
- dynamic modeling
- video compression
Fingerprint
Dive into the research topics of 'LNeRV: Disentangled Latent Representation for Semantic-Aware Video Dynamics Modeling'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver