TY - GEN
T1 - Wide-scale Building Vision Transformer used for Building Detection in Remote Sensing Images
AU - Gao, Tong
AU - Liu, Chang
AU - Wang, Shuai
AU - Si, Lingyu
AU - Dong, Hongwei
AU - Sun, Yongqiang
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - To detect buildings with diverse structures in high-resolution optical remote sensing images, the Wide-scale Building Vision Transformer (W-BViT) model is proposed by optimizing feature representation hyperparameters via analyzing spatial distribution and morphological characteristics of buildings. To extract the details of buildings, the Spatial-Detailed Context branch is constructed to learn the deep semantical information of local objects. To excavate the building correlation features, the global context branch is generated by applying vision encoding. Compared to conventional ViT, W-BViT improves the BuildFormer dual-path ViT model by constructing building-oriented convolution kernel parameters. Experiment results show that the IoU index of the W-BViT model on the WHU (WHU Building Dataset) and Massachusetts datasets is 0.4% and 0.3% higher than that of the BuildFormer model, respectively, and obtains the best detection results among all the comparison methods.
AB - To detect buildings with diverse structures in high-resolution optical remote sensing images, the Wide-scale Building Vision Transformer (W-BViT) model is proposed by optimizing feature representation hyperparameters via analyzing spatial distribution and morphological characteristics of buildings. To extract the details of buildings, the Spatial-Detailed Context branch is constructed to learn the deep semantical information of local objects. To excavate the building correlation features, the global context branch is generated by applying vision encoding. Compared to conventional ViT, W-BViT improves the BuildFormer dual-path ViT model by constructing building-oriented convolution kernel parameters. Experiment results show that the IoU index of the W-BViT model on the WHU (WHU Building Dataset) and Massachusetts datasets is 0.4% and 0.3% higher than that of the BuildFormer model, respectively, and obtains the best detection results among all the comparison methods.
KW - Building detection
KW - Object classification
KW - Remote sensing image
KW - Vision Transformer
UR - https://www.scopus.com/pages/publications/105028826317
U2 - 10.1109/EIT67313.2025.11232213
DO - 10.1109/EIT67313.2025.11232213
M3 - 会议稿件
AN - SCOPUS:105028826317
T3 - 2025 4th International Conference on Electronic Information Technology, EIT 2025
SP - 900
EP - 906
BT - 2025 4th International Conference on Electronic Information Technology, EIT 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 4th International Conference on Electronic Information Technology, EIT 2025
Y2 - 22 August 2025 through 24 August 2025
ER -