To address long-tail distribution and information loss in small traffic sign detection, STF-YOLO was proposed. A small-object conditional diffusion model was employed for augmentation on TT100K. This process created a balanced TT100K-P dataset to alleviate the long-tail issue. An Adaptive Asymptotic Feature Pyramid Network with a Multi-Channel Adaptive Attention (MC-ACA) module was then introduced. This module fuses multi-level features to recover fine-grained details for small objects. Additionally, the K-NCIoU regression loss was designed, combining the advantages of CIoU and NWD to further enhance the model's detection performance. The testing results on three traffic sign detection datasets indicate that, compared to the baseline model YOLOv8, STF-YOLO achieves improvements in mAP@0.5:0.95 of 2.5%, 3.3%, and 2.8% on the TT100K-P, TT100K, and CCTSDB datasets, respectively. Real vehicle validation confirms its excellent real-time detection capability for traffic signs.
为了确保输出图像的小目标清晰,使用特征线性调制机制(Feature-wise Linear Modulation, FiLM)[17]得到类别掩码嵌入为目标类别和位置提供参考信息。首先,如式(4)所示,类别图像掩码通过投影矩阵将其转换为类别图像特征嵌入。接着,如式(5)(6)所示,边界框嵌入通过FiLM层生成两个可学习参数与对类别图像特征嵌入进行缩放和偏移,以此得到类别掩码信息嵌入。计算过程如下:
实验采用平均精确率(mAP)、精确率(AP)、浮点运算数(FLOPs)和参数量(Params)作为检测性能评估指标。mAP@0.5表示IoU阈值为0.5时的平均精确率,mAP@0.75表示IoU阈值为0.75时的平均精确率,mAP@0.5:0.95表示IoU从0.5到0.95(步长为0.05)的平均精确率,三者直观反映模型性能。APs、APm和APl分别代表检测小型、中型和大型物体的检测精度。浮点运算次数(Floating point operations,FLOPs)用于衡量模型的复杂程度。参数量(Parms)表示模型的参数数量,衡量对显卡性能的需求。
LiuY, PengJ, XueJ H, et al. TSingNet: scale-aware and context-rich feature learning for traffic sign detection and recognition in the wild[J]. Neurocomputing, 2021, 447: 10-22.
ZhangMu-yi. Research on detection and recognition of traffic signs in complex backgrounds[D]. Xi'an: School of Aerospace Science and Techndagy, Xidian University, 2020.
[4]
YuX, LiF, LiuY, et al. A multi-stage adaptive copy-paste data augmentation algorithm based on model training preferences[J].Electronics, 2023, 12(17): 3695-3702.
[5]
ZhangT, ZouJ, JiaW. Fast and robust road sign detection in driver assistance systems[J]. Applied Intelligence, 2018, 48: 4113-4127.
[6]
YuL, XiaX, ZhouK. Traffic sign detection based on visual co-saliency in complex scenes[J]. Applied Intelligence, 2019, 49: 764-790.
[7]
BerkayaS K, GunduzH, OzsenO, et al. On circular traffic sign detection and recognition[J]. Expert Systems with Applications, 2016, 48: 67-75.
[8]
HeK, ZhangX, RenS, et al. Deep residual learning for image recognition[C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 2016: 770-778.
[9]
WuJ, LiaoS. Traffic sign detection based on SSD combined with receptive field module and path aggregation network[J].Computational Intelligence and Neuroscience, 2022, 2022(1): 4285436.
[10]
LiuS, QiL, QinH, et al. Path aggregation network for instance segmentation[C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA,2018: 8759-8768.
[11]
LiuW, AnguelovD, ErhanD, et al. Ssd: single shot multibox detector[C]∥Computer Vision-ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, 2016: 21-37.
[12]
DuS, PanW, LiN, et al. TSD‐YOLO: Small traffic sign detection based on improved YOLO v8[J]. IET Image Processing,2024, 18(11): 2884-2898.
[13]
WangX, TianY, ZhengK, et al. C2Net-YOLOv5: A bidirectional Res2Net-based traffic sign detection algorithm[J]. Computers, Materials & Continua, 2023, 77(2): 1949-1965.
[14]
YunS, HanD, OhS J, et al. CutMix: regularization strategy to train strong classifiers with localizable features[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision,2019: 6023-6032.
[15]
DvornikN, MairalJ, SchmidC. Modeling visual context is key to augmenting object detection datasets[C]∥Proceedings of the European Conference on Computer Vision (ECCV),Munich, Germany, 2018: 364-380.
[16]
GeorgakisG, MousavianA, BergA C, et al. Synthesizing training data for object detection in indoor scenes[J]. arXiv preprint arXiv:2017.
[17]
YunW H, KimT, LeeJ, et al. Cut-and-paste dataset generation for balancing domain gaps in object instance detection[J]. IEEE Access, 2021, 9: 14319-14329.
[18]
PerezE, StrubF, de VriesH, et al. FiLm: visual reasoning with a general conditioning layer[C]∥Proceedings of the AAAI Conference on Artificial Intelligence,New Orleans, Louisiana, USA, 2018: 3942-3951.
[19]
ZhengG, ZhouX, LiX, et al. LayoutDiffusion: Controllable diffusion model for layout-to-image generation[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Vancouver, BC, Canada, 2023: 22490-22499.
[20]
ZhuZ, LiangD, ZhangS, et al. Traffic sign detection and classification in the wild[C]∥Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Las Vegas, NV, USA, 2016: 2110-2118.
[21]
WangQ, WuB, ZhuP, et al. ECA-Net: Efficient channel attention for deep convolutional neural networks[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). Seattle, WA, USA, 2020: 11534-11542.
[22]
YangG, LeiJ, ZhuZ, et al. AFPN: asymptotic feature pyramid network for object detection[C]∥2023 IEEE International Conference on Systems, Man, and Cybernetics (SMC), IEEE, 2023: 2184-2189.
[23]
ZhengZ, WangP, LiuW, et al. Distance-IoU loss: Faster and better learning for bounding box regression[C]∥Proceedings of the AAAI Conference on Artificial Intelligence,New York, USA, 2020: 12993-13000.
[24]
WangJ, XuC, YangW, et al. A normalized Gaussian Wasserstein distance for tiny object detection[J]. arXiv preprint arXiv:2021.
[25]
ZhangJ, ZouX, KuangL D, et al. CCTSDB 2021: a more comprehensive traffic sign detection benchmark[J]. Human-centric Computing and Information Sciences, 2022, 12: 23.
[26]
LiX, XieZ, DengX, et al. Traffic sign detection based on improved faster R-CNN for autonomous driving[J]. The Journal of Supercomputing, 2022,78: 1-21.
[27]
LinT Y, GoyalP, GirshickR, et al. Focal loss for dense object detection[C] //Proceedings of the IEEE International Conference on Computer Vision (ICCV). Venice, Italy, 2017: 2980-2988.
[28]
CaiZ, VasconcelosN. Cascade R-CNN: Delving into high quality object detection[C] //Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Salt Lake City, UT, USA, 2018: 6154-6162.
HuJ, WangZ, ChangM, et al. Psg-yolov5: a paradigm for traffic sign detection and recognition algorithm based on deep learning[J]. Symmetry, 2022, 14(11): 2262.
[31]
YanB, LiJ, YangZ, et al. AIE-YOLO: auxiliary information enhanced YOLO for small object detection[J]. Sensors, 2022, 22(21): 8221.
[32]
ZengG, HuangW, WangY, et al. Transformer fusion and residual learning group classifier loss for long-tailed traffic sign detection[J]. IEEE Sensors Journal, 2024, 24(7): 10551-10560.