Abstract
Image sequence interpolation is a critical research area in computer vision with broad applications in video frame interpolation and medical image interlayer interpolation. Traditional deep learning-based methods in this domain predominantly rely on deep convolutional neural networks (CNNs), which, despite their effectiveness, are limited by the inherent constraints of CNN architecture, impacting their interpolation accuracy. To address these limitations, we introduce the Pre-ISIformer, a parallel multi-channel adaptive image sequence interpolation network founded on pre-trained transformers. This innovative network is composed of three integral modules: 1) Global feature extraction module is designed to extract primary features from the input images using a pre-trained Swin-transformer model, ensuring comprehensive global feature coverage. 2) Feature sequence construction module adaptively decomposes the object's motion path across different frames, facilitating a detailed analysis of motion dynamics. And 3) Intermediate image reconstruction module is responsible for accurately capturing target displacements. Furthermore, we incorporate distinct metrics for pixel loss and gradient loss to meticulously reconstruct the texture and contours of the intermediate images. Our network has been rigorously tested on various datasets for two primary applications: video frame interpolation and interlayer interpolation in medical imaging. The results from these experiments showcase the superior performance and effectiveness of the Pre-ISIformer, establishing it as a significant advancement in the field of image sequence interpolation.
| Original language | English |
|---|---|
| Pages (from-to) | 10464-10478 |
| Number of pages | 15 |
| Journal | IEEE Transactions on Circuits and Systems for Video Technology |
| Volume | 34 |
| Issue number | 10 |
| DOIs | |
| State | Published - 2024 |
| Externally published | Yes |
Keywords
- Image sequence interpolation
- interlayer interpolation
- parallel convolutional blocks
- pretrained transformers
Fingerprint
Dive into the research topics of 'Pre-Trained Transformer-Based Parallel Multi-Channel Adaptive Image Sequence Interpolation Network'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver