To solve the problem of large computation and low accuracy of the current mainstream algorithms for small object detection, this paper replaces the backbone network in YOLOv4 with the lightweight network MobileNetV3, and replaces some ordinary convolutions in the neck network with depthwise separable convolutions. At the same time, a new loss function IF-EIoU Loss is defined for small object detection. Therefore, MDS-YOLO object detection model is constructed. This model has a high detection speed and good detection performance for small object. To verify the effectiveness of the model, experiments are carried out on MS COCO dataset and Visdrone2019 dataset, respectively. Compared with the YOLOv4 algorithm, on MS COCO dataset, the average detection accuracy of the MDS-YOLO algorithm is improved by 1.5 percentage points, the detection accuracy of small object is increased by 3.3 percentage points, and the detection speed is also increased from 31 frames per second to 36 frames per second. On the Visdrone2019 dataset, the MDS-YOLO algorithm increases the average detection accuracy from 14.9% of YOLOv4 to 16.3%. The experimental results show that the MDS-YOLO algorithm proposed can effectively improve the detection accuracy of small object.
为提升小目标检测性能的有效性,本文从输入图像、网络结构、损失函数三方面改进YOLOv4算法,提出了MDS-YOLO模型.通过实验验证,在 MS COCO数据集上MDS-YOLO模型对于小目标的平均检测精度APS从YOLOv4的24.3%上升至27.6%,提升了3.3个百分点,且检测速度也提升了约16%;在Visdrone2019无人机数据集上,与YOLOv4相比,MDS-YOLO模型以牺牲较少的检测速度为代价,使检测精度提升了1.4个百分点.由此可见,本文模型的小目标检测效果更好、更高效.
ZHANGW, ZHUANGX T, WANGX L,et al .DS-YOLO:a real-time small object detection algorithm on UAVs[J].Journal of Nanjing University of Posts and Telecommunications (Natural Science Edition),2021,41(1):86-98.(in Chinese)
YAOT, YUX Y, WANGY,et al .Improvement of small target recognition algorithm of aerial photography images based on SSD[J].Ship Electronic Engineering,2020,40(9):162-166.(in Chinese)
[6]
SINGHB, DAVISL S .An analysis of scale invariance in object detection - SNIP[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City,UT,USA. IEEE,2018:3578-3587.
[7]
LINT Y, GOYALP, GIRSHICKR,et al .Focal loss for dense object detection[C]//2017 IEEE International Conference on Computer Vision (ICCV). Venice,Italy. IEEE,2017:2999-3007.
[8]
DUANK W, BAIS, XIEL X,et al .CenterNet:keypoint triplets for object detection[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). Seoul,Korea (South). IEEE,2019:6568-6577.
[9]
TIANZ, SHENC H, CHENH,et al .FCOS:fully convolutional one-stage object detection[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). Seoul,Korea (South). IEEE,2019:9626-9635.
[10]
TANM X, PANGR M, LEQ V .EfficientDet:scalable and efficient object detection[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Seattle,WA,USA. IEEE,2020:10778-10787.
LIANGX L, WANGX Y, HUANGY J, et al. Two-structure pyramid scene parsing network with attention module[J]. Journal of Changsha University of Science & Technology (Natural Science), 2024, 21(5): 104-112.(in Chinese)
[13]
BOCHKOVSKIYA, WANGC Y, LIAOH Y M .YOLOv4:optimal speed and accuracy of object detection[EB/OL].2020:2004.10934.
[14]
HOWARDA, SANDLERM, CHENB,et al .Searching for MobileNetV3[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). Seoul,Korea (South). IEEE,2019:1314-1324.
[15]
KRIZHEVSKYA, SUTSKEVERI, HINTONG E .ImageNet classification with deep convolutional neural networks[J].Communications of the ACM,2012,60:84-90.
[16]
GIRSHICKR. Fast R-CNN[C]//2015 IEEE International Conference on Computer Vision (ICCV). Santiago, Chile. IEEE,2015: 1440-1448.
[17]
CHOLLETF .Xception:deep learning with depthwise separable convolutions[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu,HI,USA. IEEE,2017:1800-1807.
[18]
SANDLERM, HOWARDA, ZHUM L,et al .MobileNetV2:inverted residuals and linear bottlenecks[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City,UT,USA. IEEE,2018:4510-4520.
[19]
LINT Y, DOLLÁRP, GIRSHICKR,et al .Feature pyramid networks for object detection[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu, HI,USA. IEEE,2017:936-944.
[20]
ZHENGZ H, WANGP, LIUW,et al .Distance-IoU loss:faster and better learning for bounding box regression[EB/OL].2019:1911.08287.
[21]
ZHANGY F, RENW Q, ZHANGZ,et al. Focal and efficient IOU loss for accurate bounding box regression[J]. Neurocomputing,2022, 506: 146-157.
[22]
CHENC Y, LIUM Y, TUZELO,et al. R-CNN for small object detection[M]//Lecture Notes in Computer Science.Cham:Springer International Publishing,2017:214-230.
[23]
RENS Q, HEK M, GIRSHICKR,et al .Faster R-CNN:towards real-time object detection with region proposal networks[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2017,39(6):1137-1149.
[24]
HEK M, ZHANGX Y, RENS Q,et al .Deep residual learning for image recognition[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas,NV,USA. IEEE,2016:770-778.
[25]
HEK M, GKIOXARIG, DOLLÁRP,et al .Mask R-CNN[C]//2017 IEEE International Conference on Computer Vision (ICCV). Venice,Italy. IEEE,2017:2980-2988.
[26]
LIUS, QIL, QINH F,et al .Path aggregation network for instance segmentation[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City,UT,USA. IEEE,2018:8759-8768.
[27]
LIUW, ANGUELOVD, ERHAND,et al .SSD:single shot MultiBox detector[M]//Lecture Notes in Computer Science.Cham:Springer International Publishing,2016:21-37.
[28]
FARHADIA, REDMONJ. YOLOv3: an incremental improvement[EB/OL]. 2018: 1804.02767.
[29]
CARIONN, MASSAF, SYNNAEVEG,et al .End-to-end object detection with transformers[M]//Lecture Notes in Computer Science.Cham:Springer International Publishing,2020:213-229.
[30]
WANGY M, ZHANGX Y, YANGT,et al .Anchor DETR: query design for transformer-based detector[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2022, 36(3): 2567-2575.
[31]
MENGD P, CHENX K, FANZ J,et al .Conditional DETR for fast training convergence[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV). Montreal,QC,Canada. IEEE,2021:3631-3640.
[32]
LIF, ZHANGH, LIUS L,et al .DN-DETR:accelerate DETR training by introducing query DeNoising[C]//IEEE Transactions on Pattern Analysis and Machine Intelligence.IEEE,2024:2239-2251.
[33]
WANGC Y, BOCHKOVSKIYA, LIAOH Y M .YOLOv7:trainable bag-of-freebies sets new state-of-the-art for real-time object detectors[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver,BC,Canada. IEEE,2023:7464-7475.
基金资助
国家自然科学基金重点资助项目(51839002)
国家自然科学基金重点资助项目(52338009)
National Natural ScienceFoundation of China(51839002)
National Natural ScienceFoundation of China(52338009)
国家杰出青年科学基金资助项目(52025085)
National Science Fund for Distinguished Young(52025085)
湖南省自然科学基金资助项目(2021JJ30734)
Natural Science Foundation of Hunan Province(2021JJ30734)
湖南省研究生创新性课题(CX20220952)
Hunan Provincial Innovation Foundation for Postgraduate(CX20220952)