Skip to main navigation Skip to search Skip to main content

Arbitrary-Scale Video Super-Resolution with Structural and Textural Priors

  • Wei Shang
  • , Dongwei Ren*
  • , Wanying Zhang
  • , Yuming Fang
  • , Wangmeng Zuo
  • , Kede Ma
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • City University of Hong Kong
  • Jiangxi University of Finance and Economics

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Arbitrary-scale video super-resolution (AVSR) aims to enhance the resolution of video frames, potentially at various scaling factors, which presents several challenges regarding spatial detail reproduction, temporal consistency, and computational complexity. In this paper, we first describe a strong baseline for AVSR by putting together three variants of elementary building blocks: 1) a flow-guided recurrent unit that aggregates spatiotemporal information from previous frames, 2) a flow-refined cross-attention unit that selects spatiotemporal information from future frames, and 3) a hyper-upsampling unit that generates scale-aware and content-independent upsampling kernels. We then introduce ST-AVSR by equipping our baseline with a multi-scale structural and textural prior computed from the pre-trained VGG network. This prior has proven effective in discriminating structure and texture across different locations and scales, which is beneficial for AVSR. Comprehensive experiments show that ST-AVSR significantly improves super-resolution quality, generalization ability, and inference speed over the state-of-the-art. The code is available at https://github.com/shangwei5/ST-AVSR.

Original languageEnglish
Title of host publicationComputer Vision – ECCV 2024 - 18th European Conference, Proceedings
EditorsAleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, Gül Varol
PublisherSpringer Science and Business Media Deutschland GmbH
Pages73-90
Number of pages18
ISBN (Print)9783031729973
DOIs
StatePublished - 2025
Event18th European Conference on Computer Vision, ECCV 2024 - Milan, Italy
Duration: 29 Sep 20244 Oct 2024

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume15115 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference18th European Conference on Computer Vision, ECCV 2024
Country/TerritoryItaly
CityMilan
Period29/09/244/10/24

Keywords

  • Arbitrary-scale video super-resolution
  • Structural and textural priors

Fingerprint

Dive into the research topics of 'Arbitrary-Scale Video Super-Resolution with Structural and Textural Priors'. Together they form a unique fingerprint.

Cite this