1.School of Electronic and Electrical Engineering,Shanghai University of Engineering Science,Shanghai 201620,China
2.Wanfeng Technology Development Co. ,Ltd,Shaoxing 312000,Zhejiang,China
3.Engineering Research Center of Intelligent Technology for Agricultural,Ministry of Education,Huazhong Agricultural University,Wuhan 430070,Hubei,China
To address the problems of slow speed, low accuracy, and insufficient inference ability in the detection of bad appearance under chip dispensing motion, the original YOLOv8n model was improved. Based on the improved YOLOv8n, a detection algorithm called Self-Position Attention-Knowledge Graph (SPA-KG), which fuses knowledge graphs and improves YOLOv8n, was proposed. First, SPA attention is designed based on coordinate attention (CA), and SPA attention can learn detailed information about small targets more fully. Second, lightweight convolution modules GHOSTConv and Adaptive Kernel Convolution (AKConv) were introduced into the backbone network and feature fusion networks, respectively, and spatial pyramid pooling was improved by using Simplified Spatial Pyramid Pooling-Fast (SIMSPPF). The number of parameters in the model was reduced to improve the detection speed, and the α-EIoU loss function was design by combining the α-IoU and EIoU to improve the localization ability and recognition accuracy of the algorithm. The experimental results showed that the average precision of the SPA-KG reached 96.7%, the precision rate reached 94.2%, and the recall rate reached 94.0%. The number of parameters reached 2.58×106, and the detection speed reached 107.7 frames/s. SPA-KG meets the industrial detection requirements.
GIRSHICKR, DONAHUEJ, DARRELLT, et al. Rich feature hierarchies for accurate object detection and semantic segmentation[C]//Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2014: 580-587. DOI: 10.1109/CVPR.2014.81 .
[2]
GIRSHICKR. Fast R-CNN[C]//2015 IEEE International Conference on Computer Vision (ICCV). New York: IEEE Press, 2015: 1440-1448. DOI: 10.1109/ICCV.2015.169 .
[3]
RENS Q, HEK M, GIRSHICKR, et al. Faster R-CNN: Towards real-time object detection with region proposal networks[C]//IEEE Transactions on Pattern Analysis and Machine Intelligence. New York: IEEE Press, 2017, 39(6): 1137-1149. DOI: 10.1109/TPAMI.2016.2577031 .
[4]
HEK M, GKIOXARIG, DOLLÁRP, et al. Mask R-CNN[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 42(2): 386-397. DOI: 10.1109/TPAMI.2018.2844175 .
[5]
LIUW, ANGUELOVD, ERHAND, et al. SSD: Single shot multibox detector[M]//Computer Vision — ECCV 2016. Cham: Springer International Publishing, 2016: 21-37. DOI: 10.1007/978-3-319-46448-0_2 .
[6]
REDMONJ, DIVVALAS, GIRSHICKR, et al. You only look once: Unified, real-time object detection[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2016: 779-788. DOI: 10.1109/CVPR.2016.91 .
[7]
REDMONJ, FARHADIA. YOLO9000: Better, faster, stronger[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2017: 6517-6525. DOI: 10.1109/CVPR.2017.690 .
LIUY J, Yilihamu·Yaermaimaiti, XIL F, et al. Research on improved safety helmet wearing detection algorithm of YOLOv5s[J]. Computer Engineering and Applications,2023,59(20):184-191 (Ch).
[13]
CHENX L, LIL J, LIF F, et al. Iterative visual reasoning beyond convolutions[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 7239-7248. DOI: 10.1109/CVPR.2018.00756 .
[14]
CHENZ M, WEIX S, WANGP, et al. Multi-label image recognition with graph convolutional networks[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2019: 5172-5181. DOI: 10.1109/cvpr.2019.00532 .
[15]
XUH, JIANGC H, LIANGX D, et al. Reasoning-RCNN: Unifying adaptive global reasoning into large-scale object detection[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2019: 6412-6421. DOI: 10.1109/cvpr.2019.00658 .
[16]
WEIW, CHENGY, HEJ F, et al. A review of small object detection based on deep learning[J]. Neural Computing and Applications, 2024, 36(12): 6283-6303. DOI: 10.1007/s00521-024-09422-6 .
[17]
YANGC, HUANGZ H, WANGN Y. QueryDet: Cascaded sparse query for accelerating high-resolution small object detection[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2022: 13658-13667. DOI: 10.1109/CVPR52688.2022.01330 .
ZHOUS Y, XUH Y, ZHUX Z, et al. Mobile phone screen defect detection algorithm based on improved YOLOv8n: PGS-YOLO [J].Computer Engineering, 2025,51(5):326-339. DOI:10.19678/J.ISSN.1000.3428.0069259(Ch ).
[20]
WANGJ Q, CHENK, YANGS, et al. Region proposal by guided anchoring[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2019: 2960-2969. DOI: 10.1109/CVPR.2019.00308 .
[21]
ZHUC C, HEY H, SAVVIDESM. Feature selective anchor-free module for single-shot object detection[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2019: 840-849. DOI: 10.1109/cvpr.2019.00093 .
[22]
TIANZ, SHENC H, CHENH, et al. FCOS: Fully convolutional one-stage object detection[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE Press, 2019: 9626-9635. DOI: 10.1109/iccv.2019.00972 .
LIUW, LIAOS C, RENW Q, et al. High-level semantic feature detection: A new perspective for pedestrian detection[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2019: 5182-5191. DOI: 10.1109/cvpr.2019.00533 .
BIANP C, ZHENGZ L, LIM L, et al. Attention fusion network based video super-resolution reconstruction[J]. Journal of Computer Applications,2021, 41 (4): 1012-1019 (Ch).
[27]
HOUQ B, ZHOUD Q, FENGJ S. Coordinate attention for efficient mobile network design[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2021: 13708-13717.. DOI: 10.1109/cvpr46437.2021.01350 .
[28]
HANK, WANGY H, TIANQ, et al. GhostNet: More features from cheap operations[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2020: 1577-1586. DOI: 10.1109/CVPR42600.2020.00165 .
[29]
DAIJ F, QIH Z, XIONGY W, et al. Deformable convolutional networks[C]//2017 IEEE International Conference on Computer Vision (ICCV).New York: IEEE Press, 2017: 764-773. DOI: 10.1109/ICCV.2017.89 .
DUM H, WUL H, SUZ. Research on deep learning based lightweight license plate detection algorithm[J]. Video Engineering, 2024, 48(3): 50-54. DOI: 10.16280/j.videoe.2024.03.012(Ch ).
SUNJ, QIANL, ZHUW D,et al. Apple detection in complex orchard environment based on improved RetinaNet[J]. Transactions of the Chinese Society of Agricultural Engineering(Transactions of the CSAE),2022,38(15):314-322 (Ch).
[36]
ZHENGZ H, WANGP, LIUW, et al. Distance-IoU loss: Faster and better learning for bounding box regression[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2020, 34(7): 12993-13000. DOI: 10.1609/aaai.v34i07.6999 .
[37]
JIANGB R, LUOR X, MAOJ Y, et al. Acquisition of localization confidence for accurate object detection[C]//Computer Vision — ECCV 2018. Cham: Springer International Publishing, 2018: 816-832. DOI: 10.1007/978-3-030-01264-9_48 .