Abstract
Recently, learned video compression (LVC) has achieved remarkable progress. Accurate temporal prior modeling is crucial for the rate-distortion (RD) performance of LVC. However, existing methods often suffer from imprecise temporal priors due to limited reconstruction ability of frame decoders and insufficient exploitation of temporal information from reconstructed frames and features. To address these issues, we first design an enhanced feature reconstructor (EFR) by integrating content-aware depthwise separable convolution (CADSC) that excels at local modeling with efficient linear attention duality (ELAD) that facilitates complete global modeling. Based on the EFR, an enhanced frame decoder (EFD) is constructed to improve transform capability, thereby generating high-quality reconstructed frames and features. Furthermore, to fully mine temporal information from these high-quality reconstructions, we propose a multi-type temporal prior mixer (MTPM). Specifically, the MTPM fuses short-term temporal information obtained via optical-flow-based feature warping with long-term temporal information captured by a proposed state-update-based method, thereby further improving the accuracy of temporal prior modeling. The temporal prior produced by the MTPM not only provides high-quality conditional guidance for frame encoding and decoding, but is also combined with the hyper prior, spatial prior, and the discrete latent representation of the previous frame to construct a diverse prior fusion entropy model (DPFEM), which further enhances entropy estimation accuracy. Experimental results demonstrate that the proposed method outperforms existing LVC models and achieves an average performance gain of 29.64% compared with the LDB configuration of the H.266/VVC reference software VTM20.2.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Broadcasting |
| DOIs | |
| State | Accepted/In press - 2026 |
| Externally published | Yes |
Keywords
- Learned video compression
- entropy model
- temporal prior modeling
Fingerprint
Dive into the research topics of 'Accurate Temporal Prior Modeling for Learned Video Compression'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver