In order to improve the accuracy and efficiency of small target detection, this paper proposed an improved detection method based on YOLOv8 to address the shortcomings of existing algorithms in small target recognition. Based on YOLOv8, this method integrated SPD (Space-to-Depth) module, which effectively avoided the information loss caused by traditional strided convolution and pooling operation. At the same time, an improvement of Fractional Fourier Transform Convolution (FT_Conv) was proposed to improve the detection accuracy and computational efficiency of the model for small targets. In addition, the C2f_BiLevel Routing Attention mechanism was used to realize dynamic sparse attention, optimize the feature fusion and object detection performance, and further improve the recognition ability of the model for small targets. Finally, the Powerful-IoU loss function was introduced to improve the area expansion of the anchor frame of the existing and enhance the focusing ability of the anchor frame. The experimental results show that compared with the original YOLOv8 model, the average accuracy (AP) of the improved model in small target detection tasks is increased by 3.29 percentage points, and the false detection and missed detection rates are significantly reduced. These results confirm that the improved YOLOv8 model has obvious performance advantages in the field of small target detection.
RENS, HEK, GIRSHICKR, et al. Faster R-CNN: Towards real-time object detection with region proposal networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(6): 1137-1149.
[2]
REDMONJ, DIVVALAS, GIRSHICKR, et al. You only look once: Unified, real-time object detection[DB/OL]. (2015-05-09)[2025-01-06].
[3]
YANGJ, WANGT. Small object detection model for remote sensing images combining super-resolution assisted reasoning and dynamic feature fusion[J]. Journal of Applied Remote Sensing, 2024, 18(2): 028503.
[4]
YINQ, HUQ, LIUH, et al. Detecting and tracking small and dense moving objects in satellite videos: A benchmark[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 5612518.
[5]
ZHANGY, ZHAOH, DUANZ, et al. Congested crowd counting via adaptive multi-scale context learning[J]. Sensors, 2021, 21(11): 3777.
[6]
JIANGH, DIAOZ, SHIT, et al. A review of deep learning-based multiple-lesion recognition from medical images: Classification, detection and segmentation[J]. Computers in Biology and Medicine, 2023, 157: 106726.
[7]
KISANTALM, WOJNAZ, MURAWSKIJ, et al. Augmentation for small object detection[DB/OL]. (2019-02-19)[2025-01-06].
[8]
AKYONF C, ONUR ALTINUCS, TEMIZELA. Slicing aided hyper inference and fine-tuning for small object detection[C]//2022 IEEE International Conference on Image Processing (ICIP), 2022: 966-970.
[9]
LIUM, JIAOL, LIUX, et al. Multi-scale contourlet knowledge guide learning segmentation[J]. IEEE Transactions on Multimedia, 2024, 26: 4831-4845.
[10]
LINT Y, DOLLÁRP, GIRSHICKR, et al. Feature pyramid networks for object detection[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017: 936-944.
[11]
CAOJ, CHENQ, GUOJ, et al. Attention-guided context feature pyramid network for object detection[DB/OL]. (2020-05-23)[2025-01-06].
[12]
JIANGY, TANZ, WANNGJ, et al. GiraffeDet: A heavy-neck paradigm for object detection[DB/OL]. (2022-02-09)[2025-01-06].
[13]
LIJ, LIANGX, WEIY, et al. Perceptual generative adversarial networks for small object detection[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017: 1951-1959.
[14]
CHENY, ZHANGC, CHENB, et al. Accurate leukocyte detection based on deformable-DETR and multi-level feature fusion for aiding diagnosis of blood diseases[J]. Computers in Biology and Medicine, 2024, 170: 107917.
[15]
XUG, LIAOW, ZHANGX, et al. Haar wavelet downsampling: A simple but effective downsampling module for semantic segmentation[J]. Pattern Recognition, 2023, 143: 109819.
[16]
SUNKARAR, LUOT. No more strided convolutions or pooling: a new CNN building block for low-resolution images and small objects[DB/OL]. (2022-08-07)[2025-01-06].
[17]
ZHUL, WANGX, KEZ, et al. BiFormer: Vision transformer with bi-level routing attention[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023: 10323-10333.
LIUC, WANGK, LIQ, et al. Powerful-IoU: More straightforward and faster bounding box regression loss with a nonmonotonic focusing mechanism[J]. Neural Networks, 2024, 170: 276-284.
[20]
ZHANGH, XUC, ZHANGS J. Inner-IoU: More effective intersection over union loss with auxiliary bounding box[DB/OL]. (2023-11-14)[2025-01-06].
[21]
ZHANGH, ZHANGS J. Shape-IoU: More accurate metric considering bounding box shape and scale[DB/OL]. (2023-12-29)[2025-01-06].