TY - GEN
T1 - Frequency-Aware Spatiotemporal Modeling with Cross-Frequency Attention for Echocardiography Left Ventricle Segmentation
AU - Li, Xiaodi
AU - Li, Hongxu
AU - Hu, Sining
AU - Shao, Shuangtong
AU - Gong, Chaoguang
AU - Chen, Zhaolin
AU - Hu, Yue
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Accurate segmentation of the left ventricle (LV) in echocardiography videos is essential for cardiac function assessment but remains challenging due to complex motion patterns, temporal dependencies, and noisy artifacts. In this work, we propose a novel frequency-aware spatiotemporal modeling network that captures both video-level and frame-level frequency characteristics for robust LV segmentation. At the video level, the extracted spatiotemporal feature tensor is decomposed into low-frequency components that capture periodic cardiac dynamics and high-frequency components that retain fine structural details. A cross-frequency attention mechanism enables high-frequency queries and values to interact with low-frequency keys, enhancing boundary learning under the guidance of global motion consistency. At the frame level, we utilize joint spatial-frequency feature extraction in the encoder and adaptively fuse high- and low-frequency components through skip connections in the decoder, enhancing segmentation precision and structural fidelity for each frame. Experiments on the CAMUS dataset demonstrate that our method outperforms state-of-the-art approaches, effectively leveraging both temporal coherence and frame-level detail for accurate LV segmentation.
AB - Accurate segmentation of the left ventricle (LV) in echocardiography videos is essential for cardiac function assessment but remains challenging due to complex motion patterns, temporal dependencies, and noisy artifacts. In this work, we propose a novel frequency-aware spatiotemporal modeling network that captures both video-level and frame-level frequency characteristics for robust LV segmentation. At the video level, the extracted spatiotemporal feature tensor is decomposed into low-frequency components that capture periodic cardiac dynamics and high-frequency components that retain fine structural details. A cross-frequency attention mechanism enables high-frequency queries and values to interact with low-frequency keys, enhancing boundary learning under the guidance of global motion consistency. At the frame level, we utilize joint spatial-frequency feature extraction in the encoder and adaptively fuse high- and low-frequency components through skip connections in the decoder, enhancing segmentation precision and structural fidelity for each frame. Experiments on the CAMUS dataset demonstrate that our method outperforms state-of-the-art approaches, effectively leveraging both temporal coherence and frame-level detail for accurate LV segmentation.
KW - Cross-frequency attention
KW - Frequency-aware
KW - Segmentation
KW - Spatiotemporal modeling
UR - https://www.scopus.com/pages/publications/105041603206
U2 - 10.1109/ISBI61048.2026.11515709
DO - 10.1109/ISBI61048.2026.11515709
M3 - 会议稿件
AN - SCOPUS:105041603206
T3 - Proceedings - International Symposium on Biomedical Imaging
BT - ISBI 2026 - 23rd IEEE International Symposium on Biomedical Imaging
PB - IEEE Computer Society
T2 - 23rd IEEE International Symposium on Biomedical Imaging, ISBI 2026
Y2 - 8 April 2026 through 11 April 2026
ER -