TY - GEN
T1 - Flow-Guided ConvLSTM with Quality-Aware Reconstruction for Learned Video Compression
AU - Deng, Xuan
AU - Meng, Xiandong
AU - Man, Hengyu
AU - Fan, Xiaopeng
AU - Zhao, Debin
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - In Learned Video Compression (LVC), contextual information plays a crucial role in reducing temporal redundancy and improving rate-distortion performance. However, most existing approaches rely on a single reference frame or a fixed set of neighboring frames, which limits their ability to exploit long-range temporal dependencies. To overcome this limitation, we propose FGC-QAR (Flow-Guided ConvLSTM with Quality-Aware Reconstruction), a unified framework that combines a Flow-Guided ConvLSTM (FG-ConvLSTM) with a Quality-Aware Reconstruction (QAR) module. FG-ConvLSTM aggregates features across extended temporal spans by explicitly incorporating optical flow as motion guidance. Meanwhile, QAR is implemented using ConvGRU and dynamically modulates its forget and input gates according to the compression quality and content similarity of reference frames. This allows the network to emphasize more reliable temporal cues while suppressing noisy or low-quality information. Extensive experiments demonstrate that the proposed framework consistently improves compression efficiency across multiple datasets and bitrates, highlighting its effectiveness and scalability.
AB - In Learned Video Compression (LVC), contextual information plays a crucial role in reducing temporal redundancy and improving rate-distortion performance. However, most existing approaches rely on a single reference frame or a fixed set of neighboring frames, which limits their ability to exploit long-range temporal dependencies. To overcome this limitation, we propose FGC-QAR (Flow-Guided ConvLSTM with Quality-Aware Reconstruction), a unified framework that combines a Flow-Guided ConvLSTM (FG-ConvLSTM) with a Quality-Aware Reconstruction (QAR) module. FG-ConvLSTM aggregates features across extended temporal spans by explicitly incorporating optical flow as motion guidance. Meanwhile, QAR is implemented using ConvGRU and dynamically modulates its forget and input gates according to the compression quality and content similarity of reference frames. This allows the network to emphasize more reliable temporal cues while suppressing noisy or low-quality information. Extensive experiments demonstrate that the proposed framework consistently improves compression efficiency across multiple datasets and bitrates, highlighting its effectiveness and scalability.
KW - deep learning
KW - learned video compression
KW - quality-aware reconstruction
UR - https://www.scopus.com/pages/publications/105040904333
U2 - 10.1109/DCC66757.2026.00010
DO - 10.1109/DCC66757.2026.00010
M3 - 会议稿件
AN - SCOPUS:105040904333
T3 - Data Compression Conference Proceedings
SP - 23
EP - 32
BT - Proceedings - DCC 2026
A2 - Bilgin, Ali
A2 - Fowler, James E.
A2 - Serra-Sagrista, Joan
A2 - Ye, Yan
A2 - Storer, James A.
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2026 Data Compression Conference, DCC 2026
Y2 - 24 March 2026 through 27 March 2026
ER -