Skip to main navigation Skip to search Skip to main content

Pre-Trained Transformer-Based Parallel Multi-Channel Adaptive Image Sequence Interpolation Network

  • Hui Liu
  • , Gongguan Chen
  • , Meng Liu*
  • , Liqiang Nie
  • *Corresponding author for this work
  • Shandong University of Finance and Economics
  • Shandong Key Laboratory of Digital Media Technology
  • Shandong Jianzhu University
  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Image sequence interpolation is a critical research area in computer vision with broad applications in video frame interpolation and medical image interlayer interpolation. Traditional deep learning-based methods in this domain predominantly rely on deep convolutional neural networks (CNNs), which, despite their effectiveness, are limited by the inherent constraints of CNN architecture, impacting their interpolation accuracy. To address these limitations, we introduce the Pre-ISIformer, a parallel multi-channel adaptive image sequence interpolation network founded on pre-trained transformers. This innovative network is composed of three integral modules: 1) Global feature extraction module is designed to extract primary features from the input images using a pre-trained Swin-transformer model, ensuring comprehensive global feature coverage. 2) Feature sequence construction module adaptively decomposes the object's motion path across different frames, facilitating a detailed analysis of motion dynamics. And 3) Intermediate image reconstruction module is responsible for accurately capturing target displacements. Furthermore, we incorporate distinct metrics for pixel loss and gradient loss to meticulously reconstruct the texture and contours of the intermediate images. Our network has been rigorously tested on various datasets for two primary applications: video frame interpolation and interlayer interpolation in medical imaging. The results from these experiments showcase the superior performance and effectiveness of the Pre-ISIformer, establishing it as a significant advancement in the field of image sequence interpolation.

Original languageEnglish
Pages (from-to)10464-10478
Number of pages15
JournalIEEE Transactions on Circuits and Systems for Video Technology
Volume34
Issue number10
DOIs
StatePublished - 2024
Externally publishedYes

Keywords

  • Image sequence interpolation
  • interlayer interpolation
  • parallel convolutional blocks
  • pretrained transformers

Fingerprint

Dive into the research topics of 'Pre-Trained Transformer-Based Parallel Multi-Channel Adaptive Image Sequence Interpolation Network'. Together they form a unique fingerprint.

Cite this