Aiming at the problems of low detection accuracy, strong background interference and weak feature expression ability of small targets in unmanned aerial vehicle (UAV) remote sensing images, a target detection algorithm combining multi-core feature extraction and triplet attention mechanism is proposed. Based on the YOLOv12 framework, the MogaNet dual-domain aggregation mechanism is introduced to enhance the multi-scale feature extraction ability, the coordinate attention (CA) module is embedded to improve the spatial positioning accuracy, and the triple attention mechanism is designed to jointly strengthen the attention to small targets in the spatial, channel and scale dimensions, and the CMUNeXt feature fusion strategy is combined to optimize cross-level information transmission. The research results show that the index ImAP @ 0.5 of the algorithm on the VisDrone2019 and NWPU VHR-10 datasets reaches 0.331 and 0.808, respectively, which is 5.7% and 15.0% higher than the YOLOv12 model, and significantly improves the detection accuracy and robustness of small targets in complex backgrounds.
CA机制将二维特征图在水平与垂直方向上分别进行全局平均池化,编码出目标在不同方向上的精确位置信息。这种设计对于无人机视角下密集分布的小目标(如拥挤的人群、停车场车辆)尤为重要。因为在密集场景下,传统注意力机制容易将相邻目标的特征混淆,而 CA 机制引入的坐标信息能够辅助模型更好地区分相邻目标的边界,抑制冗余背景干扰,从而显著提升模型在密集小目标检测中的定位精度与鲁棒性。CA注意力机制模型架构见图3。
IPre(准确率)用于衡量预测为正样本的结果中,真实正样本的占比。IRec(召回率)用于衡量所有真实目标被正确检测出的比例。ImAP@0.5(mean average precision at IoU is 0.5)是指IoU阈值为0.5时所有类别平均精度(AP)的均值。ImAP@0.5:0.95为IoU阈值在0.50~0.95(步长为 0.05)范围内,所有类别平均精度的均值,用于评估模型的综合表现。
GUEBSIR, MAMIS, CHOKMANIK. Drones in precision agriculture: a comprehensive review of applications, technologies, and challenges[J]. Drones, 2024, 8(11): 686.
WUYiquan, TONGKang. Research advances on deep learning-based small object detection in UAV aerial images[J]. Acta Aeronautica et Astronautica Sinica,2025,46(3):174-200.
[4]
LINT Y, GOYALP, GIRSHICKR, et al. Focal loss for dense object detection[C]//2017 IEEE International Conference on Computer Vision. October 22-29, 2017, Venice, Italy. IEEE, 2017: 2999-3007.
[5]
LEEJ H, GWONG H, KIMI H, et al. A motion deblurring network for enhancing UAV image quality in bridge inspection[J]. Drones, 2023, 7(11): 657.
[6]
GAOT, NIUQ Q, ZHANGJ, et al. Global to local: a scale-aware network for remote sensing object detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 5615614.
XUEGuanghui, YANZhaoyang, WUMian, et al. PCSED-YOLO: study on cross-scale multi-object wearable detection algorithm in complex environments[J]. Computer Engineering and Applications, 2026, 62(5): 88-105.
[9]
WANGX L, CHENH. HPS-DETR: enhancing small object detection with lightweight feature extraction and transformer integration[J]. IEEE Transactions on Geoscience and Remote Sensing, 2025, 63: 5937420.
[10]
WUQ G, LIY, YINJ R, et al. LGC-YOLO: local-global feature extraction and coordination network with contextual interaction for remote sensing object detection[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18: 15376-15393.
ZHOUXiaoqin, WANGChanghai, WANGWei. A UAV aerial object detection model integrating contextual perception[J]. Science of Surveying and Mapping, 2026, 51(1): 88-97.
[13]
LIZ, WANGY C, ZHANGY X, et al. Context feature integration and balanced sampling strategy for small weak object detection in remote sensing imagery[J]. IEEE Geoscience and Remote Sensing Letters, 2024, 21: 6009105.
LIUHaonan, ZHOUGang, LIUJiangtao, et al. Cotton field pest detection based on super-resolution reconstruction and dynamic convolution[J/OL]. Computer Engineering, 2025: 1-17[2025-11-10].
SONGRunze, CUIHongbo, ZHAOZilong, et al. Small target detection method based on diffusion super-resolution reconstruction[J/OL]. Systems Engineering and Electronics,2025: 1-9[2025-11-13].
[22]
GUOM H, LUC Z, LIUZ N, et al. Visual attention network[J]. Computational Visual Media, 2023, 9(4): 733-752.
PENGJishen, MALongze, SUNMengyu, et al. A multi-scale feature fusion enhanced detection model MFFE-YOLO[J]. Journal of Liaoning Technical University(Natural Science),2024,43(5):625-632.
[25]
LIUW J, WUG Q, WANGH, et al. Dense skip-attention for convolutional networks[J]. Scientific Reports, 2025, 15: 22710.
ZHANGZaiyan, SONGWeidong, WUJiachen. Disease detection of cement pavement based on improved YOLOv5 in complex scenarios[J]. Journal of Liaoning Technical University (Natural Science),2025,44(1):102-112.
CHENMengyuan, XURuiheng, YANGSupeng, et al. SLAM algorithm based on improved YOLOv6s network in dynamic occlusion scenarios[J]. Journal of Chinese Inertial Technology,2025,33(8):802-811.
[30]
VARGHESER, M S. YOLOv8: a novel object detection algorithm with enhanced performance and robustness[C]//2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems. April 18-19, 2024, Chennai, India. IEEE, 2024: 1-6.
[31]
CHENQ Z. Traffic object detection using YOLOv12[J]. Open Access Library Journal, 2025, 12: e13991.
DUANHongyu, SUNZhiyang. A YOLOv13-based method for drone type detection[J]. Computer Era,2025(12):14-18.
[34]
RABBIJ, RAY N, SCHUBERTM, et al. Small-object detection in remote sensing images with end-to-end edge-enhanced GAN and object detector network[J]. Remote Sensing, 2020, 12(9): 1432.
WANGZongyang, HUANGLi, JIANGDu. APW-YOLOv8-based detection of small targets in high-altitude UAV image[J]. Computer Systems & Applications, 2025, 34(10): 195-205.
[37]
HUANGFUZ M, LIS Q, YANL H. Ghost-YOLO v8:an attention-guided enhanced small target detection algorithm for floating litter on water surfaces[J]. Computers, Materials & Continua,2024,80(3): 3713-3731.