TY - GEN
T1 - Attention-based Global Feature Extraction Method for Image Retrieval
AU - He, Xiangyun
AU - Ma, Lin
AU - Zhao, Weiqiang
AU - Qin, Danyang
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - Image retrieval technology has become increasingly important in modern society. Particularly, instance-level image retrieval not only enables fast and accurate image retrieval but also plays a significant role in the field of Visual place recognition (VPR). NetVLAD's (Vector of locally aggregated descriptors) global feature extraction method exhibits excellent performance in the field of image retrieval. Nevertheless, the NetVLAD method is characterized by a long training time, a low recall rate, and a slow method for extracting features. To address this problem, we propose an end-to-end model based on a novel deep neural network for global feature extraction. It exhibits lower computation complexity for high-dimensional features, accelerates feature extraction, and improves the recall rate of image retrieval. We present the following two principal contributions. First, we introduce attention aggregation module based on self-attention mechanism that combines the positional information and confidence scores of local features. And then it enhances local features using the self-attention mechanism. Second, we present the global feature extraction module, At-tnVLAD, based on the principles of the NetVLAD method. It employs a cross-attention mechanism in place of convolution, reducing the number of trainable parameters without affecting the computational complexity. The experimental results indicate that the proposed method can not only speed up training convergence and feature extraction but also improve recall rate.
AB - Image retrieval technology has become increasingly important in modern society. Particularly, instance-level image retrieval not only enables fast and accurate image retrieval but also plays a significant role in the field of Visual place recognition (VPR). NetVLAD's (Vector of locally aggregated descriptors) global feature extraction method exhibits excellent performance in the field of image retrieval. Nevertheless, the NetVLAD method is characterized by a long training time, a low recall rate, and a slow method for extracting features. To address this problem, we propose an end-to-end model based on a novel deep neural network for global feature extraction. It exhibits lower computation complexity for high-dimensional features, accelerates feature extraction, and improves the recall rate of image retrieval. We present the following two principal contributions. First, we introduce attention aggregation module based on self-attention mechanism that combines the positional information and confidence scores of local features. And then it enhances local features using the self-attention mechanism. Second, we present the global feature extraction module, At-tnVLAD, based on the principles of the NetVLAD method. It employs a cross-attention mechanism in place of convolution, reducing the number of trainable parameters without affecting the computational complexity. The experimental results indicate that the proposed method can not only speed up training convergence and feature extraction but also improve recall rate.
KW - Attention mechanism
KW - Global feature extraction
KW - Image retrieval
KW - Recall rate
KW - Visual place recognition
UR - https://www.scopus.com/pages/publications/85206163672
U2 - 10.1109/VTC2024-Spring62846.2024.10683309
DO - 10.1109/VTC2024-Spring62846.2024.10683309
M3 - 会议稿件
AN - SCOPUS:85206163672
T3 - IEEE Vehicular Technology Conference
BT - 2024 IEEE 99th Vehicular Technology Conference, VTC2024-Spring 2024 - Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 99th IEEE Vehicular Technology Conference, VTC2024-Spring 2024
Y2 - 24 June 2024 through 27 June 2024
ER -