The images collected by fish eye cameras in autonomous driving scenarios have severe distortion, complex scenes, drastic scale changes, and many small targets, which lead to low detection accuracy of traditional object detection models. Therefore, YOLOv5s-R, an improved fish eye image detection model based on YOLOv5s, is proposed. Firstly, to solve the problem of difficult recognition of minor targets, the RCMS (Random Crop Muti Scale) data augmentation method is proposed, which performs better than the optimal data augmentation method obtained from ablation experiments. Secondly, to improve the detection accuracy of the model, SA (Shuffle Attention) and LDH (Light Decouple Head) modules are added to the network header to enhance the model’s feature extraction and recognition capabilities, suppress noise interference. Finally, an additional angle prediction branch is added to realize the rotating box object detection, a circular label is constructed to solve the PoA (Periodicity of Angular) problem, and the label is smoothed with the Gaussian function. The RIOU is proposed to optimize the loss function by adding an angle penalty term on the basis of CIOU, which improves the regression accuracy and speeds up the convergence of the model. The experimental results show that the proposed YOLOv5s-R model achieves good detection performance on the Woodscape dataset. Compared to the original YOLOv5s model, mAP@0.5 mAP@0.5 is 0.95 increased by 6.8% and 5.6%, respectively, reaching 82.6% and 49.5%.
WoodScape数据集以小目标为主,大多的数据增强方法都不能很好地提升模型的检测能力,因此提出一种针对小目标检测的随机裁剪多尺度(Random Crop Multi Scale, RCMS)数据增强方法,如图4所示.将原图以不同尺度(图中不同颜色框)随机裁剪图像,再缩放到统一的尺寸用于训练.该方法起着“放大镜”的作用,能使模型增强对小物体的特征的提取能力,多尺度则可以增强模型辨别不同大小物体的能力.
MAOZ H, ZHUJ L, WUX,et al .Review of YOLO based target detection for autonomous driving[J].Computer Engineering and Applications,2022,58(15):68-77.(in Chinese)
DUANX T, ZHOUY K, TIAND X,et al .A review of deep learning applications for autonomous driving[J].Unmanned Systems Technology,2021,4(6):1-27.(in Chinese)
[9]
GIRSHICKR, DONAHUEJ, DARRELLT,et al .Rich feature hierarchies for accurate object detection and semantic segmentation[C]//Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, ACM,2014:580-587.
[10]
GIRSHICKR .Fast R-CNN[C]//2015 IEEE International Conference on Computer Vision (ICCV).Santiago,Chile: IEEE,2015:1440-1448.
[11]
RENS Q, HEK M, GIRSHICKR,et al .Faster R-CNN:towards real-time object detection with region proposal networks[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2017,39(6):1137-1149.
[12]
HEK M, GKIOXARIG, DOLLÁRP,et al .Mask R-CNN[C]//2017 IEEE International Conference on Computer Vision (ICCV).Venice,Italy: IEEE,2017:2980-2988.
[13]
REDMONJ, DIVVALAS, GIRSHICKR,et al .You only look once:unified,real-time object detection[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).Las Vegas,NV,USA: IEEE,2016:779-788.
[14]
REDMONJ, FARHADIA .YOLO9000:better,faster,stronger[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).Honolulu,HI,USA. IEEE,2017:6517-6525.
[15]
REDOMONJ, FARHADIA.YOLOv3: an incremental Improvement[J]. 2018.
[16]
BOCHKOVSKIYA, WANGC Y, LIAOH. Yolov4: optimal speed and accuracy of object detection[J]. arXiv preprint arXiv: 2020.
[17]
JOVHERG. YOLOv5[EB/OL]. 2021.
[18]
LIC Y, LIL, JIANGH L,et al .YOLOv6:a single-stage object detection framework for industrial applications[J]. arXiv preprint arXiv:2022.
[19]
WANGC Y, BOCHKOVSKIYA, LIAOH Y M.YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors[C]// arXiv preprint arXiv: 2022.
[20]
JOCHERG, CHAURASIAA, Q J. YOLO by Ultralytics (Version 8.0.0) [EB/OL]. 2023.
[21]
XUS, WANGX, LVW,et al.PP-YOLOE: An evolved version of YOLO[J]. 2022.
[22]
LIT W, TONGG J, TANGH Y,et al .FisheyeDet:a self-study and contour-based object detector in fisheye images[J].IEEE Access,2020,8:71739-71751.
[23]
KUMARV R, YOGAMANIS, RASHEDH,et al .OmniDet:surround view cameras based multi-task visual perception network for autonomous driving[J].IEEE Robotics and Automation Letters,2021,6(2):2830-2837.
[24]
RASHEDH, MOHAMEDE, SISTUG,et al .Generalized object detection on fisheye cameras for autonomous driving:dataset,representations and baseline[C]//2021 IEEE Winter Conference on Applications of Computer Vision (WACV).Waikoloa,HI,USA: IEEE,2021:2271-2279.
[25]
COORSB, CONDURACHEA P, GEIGERA .SphereNet:learning spherical representations for detection and classification in omnidirectional images[C]//Computer Vision – ECCV 2018:15th European Conference,Munich,Germany,September 8–14,2018,Proceedings,Part IX: ACM,2018:525–541.
[26]
ZHANGZ H, XUY Y, YUJ Y,et al .Saliency detection in 360 $$^\circ $$ videos[C]//European Conference on Computer Vision.Cham:Springer,2018:504-520.
[27]
SUY C, GRAUMANK .Kernel transformer networks for compact spherical convolution[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Long Beach,CA,USA: IEEE,2019:9434-9443.
[28]
DOOLEYD, MCGINLEYB, HUGHESC,et al .A blind-zone detection method using a rear-mounted fisheye camera with combination of vehicle detection methods[J].IEEE Transactions on Intelligent Transportation Systems,2016,17(1):264-278.
[29]
BAEKI, DAVIESA, YANG,et al .Real-time detection,tracking,and classification of moving and stationary objects using multiple fisheye images[C]//2018 IEEE Intelligent Vehicles Symposium (IV).Changshu,China: IEEE,2018:447-452.
QINX H, HUANGQ D, CHANGD X,et al .Object detection method in open-pit mine based on improved YOLOv5[J].Journal of Hunan University (Natural Sciences),2023,50(2): 23-30.(in Chinese)
YANGR N, HUIF, JINX,et al .Roadside target detection algorithm for complex traffic scene based on improved YOLOv5s[J].Computer Engineering and Applications,2023,59(16):159-169.(in Chinese)
LIUC X, LIC, PANL H,et al .Improved coal mine smoke and fire detection algorithm of YOLOv5s[J].Computer Engineering and Applications,2023,59(17):286-294.(in Chinese)
YANGC, SHEL, YANGL,et al .Improved YOLOv5 object detection algorithm for remote sensing images[J].Computer Engineering and Applications,2023,59(15):76-86.(in Chinese)
[40]
YOGAMANIS, HUGHESC, HORGANJ,et al .WoodScape:a multi-task,multi-camera fisheye dataset for autonomous driving[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV).Seoul,Korea (South). IEEE,2019:9307-9317.
[41]
ZHANGQ L, YANGY B .SA-net:shuffle attention for deep convolutional neural networks[C]//ICASSP 2021 - 2021 IEEE International Conference on Acoustics,Speech and Signal Processing (ICASSP).Toronto,ON,Canada. IEEE,2021:2235-2239.
[42]
WUY, CHENY P, YUANL,et al .Rethinking classification and localization for object detection[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Seattle,WA,USA: IEEE,2020:10183-10192.
[43]
GEZ, LIUS, WANGF, et al. Yolox: exceeding yolo series in 2021[J].2021.
[44]
YANGX, YANJ C .Arbitrary-oriented object detection with circular smooth label[C]//European Conference on Computer Vision.Cham:Springer,2020:677-694.