TY - GEN
T1 - Attention-enhanced YOLOv11N for traffic small object detection on VisDrone dataset
AU - Zhang, Jili
AU - Quan, Wei
AU - Liu, Chunjiang
AU - Yan, Yuchen
AU - Chen, Xuan
AU - Wang, Hua
N1 - Publisher Copyright:
© 2026 COPYRIGHT SPIE.
PY - 2026/4/29
Y1 - 2026/4/29
N2 - Traffic small object detection based on unmanned aerial vehicle (UAV) images is crucial for intelligent transportation systems. However, the VisDrone dataset, which is widely used for UAV-based detection, poses significant challenges such as small object size, dense distribution, and complex backgrounds. To address these issues, this paper proposes an optimized traffic small object detection model based on attention-enhanced YOLOv11N. First, a Dual-Branch Attention Fusion Module (DBAFM) is designed to integrate local spatial details and global semantic information, enhancing the model's ability to capture small object features. Second, an enhanced feature fusion neck structure is introduced to strengthen the propagation of low-level small object features. Finally, a coordinate-aware loss function is adopted to improve the localization accuracy of small objects. Extensive experiments are conducted on the VisDrone-2019-DET dataset. The results show that the proposed model achieves 56.8% mAP50 and 32.4% mAP50:95, which are 6.2% and 7.5% higher than the baseline YOLOv11N, respectively. Meanwhile, the model maintains a real-time inference speed of 82 FPS, demonstrating superior performance in terms of both accuracy and efficiency. It provides a reliable solution for traffic small object detection in UAV scenarios.
AB - Traffic small object detection based on unmanned aerial vehicle (UAV) images is crucial for intelligent transportation systems. However, the VisDrone dataset, which is widely used for UAV-based detection, poses significant challenges such as small object size, dense distribution, and complex backgrounds. To address these issues, this paper proposes an optimized traffic small object detection model based on attention-enhanced YOLOv11N. First, a Dual-Branch Attention Fusion Module (DBAFM) is designed to integrate local spatial details and global semantic information, enhancing the model's ability to capture small object features. Second, an enhanced feature fusion neck structure is introduced to strengthen the propagation of low-level small object features. Finally, a coordinate-aware loss function is adopted to improve the localization accuracy of small objects. Extensive experiments are conducted on the VisDrone-2019-DET dataset. The results show that the proposed model achieves 56.8% mAP50 and 32.4% mAP50:95, which are 6.2% and 7.5% higher than the baseline YOLOv11N, respectively. Meanwhile, the model maintains a real-time inference speed of 82 FPS, demonstrating superior performance in terms of both accuracy and efficiency. It provides a reliable solution for traffic small object detection in UAV scenarios.
KW - Attention Mechanism
KW - Intelligent Transportation
KW - Small Object Detection
KW - VisDrone Dataset
KW - YOLOv11N
UR - https://www.scopus.com/pages/publications/105040116904
U2 - 10.1117/12.3115649
DO - 10.1117/12.3115649
M3 - 会议稿件
AN - SCOPUS:105040116904
T3 - Proceedings of SPIE - The International Society for Optical Engineering
BT - Second International Conference on Image Processing and Deep Learning, IPDL 2026
A2 - Wang, Jun
A2 - Leng, Lu
PB - SPIE
T2 - 2nd International Conference on Image Processing and Deep Learning, IPDL 2026
Y2 - 6 March 2026 through 8 March 2026
ER -