1.Guangdong Key Laboratory of Intelligent Transportation System,School of Intelligent Systems Engineering,Sun Yat-sen University,Guangzhou,510275,China
2.Hunan Expressway Information Technology Co. ,Ltd,Changsha 410001,China
To address the challenge of distinguishing vehicles from the background under low-light conditions and the increased detection difficulty caused by light source interference, this paper proposes an improved lightweight vehicle detection model—Global Enhancement and Local Attention YOLO (GALA-YOLO). Existing deep learning-based object detection models are primarily designed for daytime scenarios, overlooking the unique challenges of nighttime environments. Vehicle detection in low-light nighttime scenarios has become a critical and challenging aspect of advanced driver assistance system (ADAS) research. The model is optimized based on the lightweight YOLOv8s, incorporating Global Attention (GA) and Local Attention (LA) modules, which are embedded at different layers to enhance feature representation. The GA module uses large convolutional kernels to capture global information, such as object spatial distribution and lighting intensity, with adaptive adjustments for nighttime environments. The LA module is a multi-scale convolutional attention module that employs convolutional kernels of varying sizes to capture multi-scale local information, considering the different levels of detail in complex nighttime environments. Furthermore, the paper improves the CIoU loss function by adopting MPDIoU, which allows for more flexible capture of boundary box shape features, significantly improving the model's robustness and convergence speed. Experimental results demonstrate that GALA-YOLO achieves an mAP of 59.2% on nighttime scene datasets, a 0.8% improvement over the best existing object detection model. On the KITTI dataset, the vehicle detection mAP50 reaches 76.7%, surpassing the current optimal detection model by 0.5%. The results validate the effectiveness and robustness of the proposed method in nighttime environments, providing a reference for intelligent detection and safety warning research inADAS(Advanced Driver Assistance Systems).
针对夜间光照条件差和特征信息匮乏等问题,本文基于轻量级的YOLOv8s模型进行了多方面的改进,旨在保证实时性的前提下,提升夜间车辆检测的鲁棒性和精度。具体贡献总结如下:①全局增强模块(Global augmentation, GA):为了应对夜间低光照条件下的低对比度问题,本文提出了全局增强模块。该模块通过采用大卷积核来捕捉图像的全局信息,增强上下文特征,并能够自适应地调整光照强度和颜色分布等参数,以有效缓解低亮度和低对比度对目标检测性能的负面影响。②局部注意力模块(Local Attention, LA):在夜间环境中,由于车辆与背景之间的差异通常较小,局部特征的提取变得尤为关键。为此,本文引入了局部注意力模块,该模块通过多尺度深度卷积进行多分支局部特征的提取与聚合,从而增强了模型对图像中不同尺度的局部信息和细节特征的敏感性。③损失函数改进:针对现有损失函数在夜间车辆检测中的不足,本文引入了 MPDIoU(Minimum point distance based IoU),该损失函数专门设计用于优化目标框的定位精度。MPDIoU损失能够有效应对目标物体在图像中发生重叠,或检测框与真实框长宽比相同但大小不同的情况,从而实现更加精确的目标定位。
IqbalA, RehmanZ U, AliS, et al. Road traffic accident analysis and identification of black spot locations on highway [J]. Civil Engineering Journal, 2020, 6(12): 2448-2456.
[2]
HangJ, YanX, LiX, et al. In-vehicle warnings for work zone and related rear-end collisions: a driving simulator experiment [J]. Accident Analysis & Prevention, 2022, 174: 106768.
[3]
ArthursP, GillamL, KrauseP, et al. A taxonomy and survey of edge cloud computing for intelligent transportation systems and connected vehicles [J]. IEEE Transactions on Intelligent Transportation Systems, 2021, 23(7): 6206-6221.
HuangLing, GuoHeng-cong, ZhangRong-hui, et al. LSTM-based lane-changing behavior model for unmanned vehicle under environment of heterogeneous human-driven and autonomous vehicles [J]. China Journal of Highway and Transport, 2020, 33(7): 156-166.
ZhangRong-hui, YouFeng, ChuXin-nan, et al. Lane change merging control method for unmanned vehicle under V2V cooperative environment[J]. China Journal of Highway and Transport, 2018, 31(4): 180-191.
[8]
ZhangX, StoryB, RajanD. Night time vehicle detection and tracking by fusing vehicle parts from multiple cameras[J]. IEEE Transactions on Intelligent Transportation Systems, 2021, 23(7): 8136-8156.
[9]
BaiP F. Nighttime vehicle detection method based on headlight features[J]. Electronic Measurement Technology, 2020, 43(14): 89-95.
[10]
KavyaT S, TsogtbaatarE, JangY M, et al. Night-time vehicle detection based on brake/tail light color[C]∥International SoC Design Conference (ISOCC), Daegu, Korea (South), 2018: 206-207.
[11]
KuangH, ChenL, ChanL L H, et al. Feature selection based on tensor decomposition and object proposal for night-time multiclass vehicle detection[J]. IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2018, 49(1): 71-80.
[12]
MoY, HanG, ZhangH, et al. Highlight-assisted nighttime vehicle detection using a multi-level fusion network and label hierarchy[J]. Neurocomputing, 2019, 355: 13-23.
FanJia-qi, LiXin, HuoTian-jiao, et al. Research on cross domain detection in intelligent vehicles based on one-stage algorithm[J]. China Journal of Highway and Transport, 2022, 35(3): 249-262.
ChengXin, WangHong-fei, ZhouJing-mei, et al. Vehicle detection algorithm based on voxel pillars from LiDAR point clouds[J]. China Journal of Highway and Transport, 2023, 36(3): 247-260.
ZhangBing-li, PanZe-hao, JiangJun-zhao, et al. Multi-modal perception fusion method based on cross attention[J]. China Journal of Highway and Transport, 2024, 37(03): 181-193.
[19]
ArinaldiA, PradanaJ A, GurusingaA A. Detection and classification of vehicles for traffic video analytics[J]. Procedia Computer Science, 2018, 144: 259-268.
[20]
ParkJ M, LeeJ W. Combination of SSD and a Rule-based approachfor nighttime vehicle detection on roads[J].Journal of Institute of Control, Robotics and Systems,2020, 26(6): 493-498.
[21]
ChenL, HuX, XuT, et al. Turn signal detection during nighttime by CNN detector and perceptual hashing tracking[J]. IEEE Transactions on Intelligent Transportation Systems, 2017, 18(12): 3303-3314.
[22]
RedmonJ. You only look once: unified, real-time object detection[C]∥IEEE Conference on Computer Vision and Pattern Recognition (CVPR),Las Vegas, Nevada, USA, 2016: 779-788.
WangAi-di, PengYi-chuan, LangHong, et al. Pavement pothole extraction based on YOLOX-transformer two-step model[J]. China Journal of Highway and Transport, 2023, 36(12): 304-317.
HuXiao-wei, Yan Yi-xin Wang Da-wei, et al. Lightweight pavement disease detection based on YOLOM algorithm[J]. China Journal of Highway and Transport, 2024, 37(12): 381-391.
DengShi-qiang, DingHao, JiangShu-ping, et al. Intelligent detection algorithm for fire smoke in highway tunnel based on improved YOLOv5s[J]. China Journal of Highway and Transport, 2024, 37(11): 194-209.
[29]
HeK, ZhangX, RenS, et al. Spatial pyramid pooling in deep convolutional networks for visual recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2015, 37(9): 1904-1916.
[30]
LinT Y, DollárP, GirshickR, et al. Feature pyramid networks for object detection[C]∥IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 2017: 2117-2125.
[31]
DaiJ, QiH, XiongY, et al. Deformable convolutional networks[C]∥2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 2017: 764-773.
[32]
DingX, ZhangX, HanJ, et al. Scaling up your kernels to 31×31: Revisiting large kernel design in CNNs[C]∥IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR),New Orleans, LA, USA, 2022: 11963-11975.
ZhangFan, WangJing-bo, ZhuYang, et al. Lightweight object detection algorithm for ship recognition in remote sensing image[J]. Journal of Jilin University (Engineering and Technology Edition),2026, 56(2): 533-542.
GaoYun-long, RenMing, WuChuan, et al. An improved anchor-free model based on attention mechanism for ship detection[J]. Journal of Jilin University (Engineering and Technology Edition), 2024, 54(5): 1407-1416.
[37]
PengC, ZhangX, YuG, et al. Large kernel matters-improve semantic segmentation by global convolutional network[C]∥IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 2017: 4353-4361.
[38]
HouQ, ZhangL, ChengM M, et al. Strip pooling: Rethinking spatial pooling for scene parsing[C]∥IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2020: 4003-4012.
[39]
ZhengZ, WangP, LiuW, et al. Distance-IoU loss: faster and better learning for bounding box regression[C]∥AAAI Conference on Artificial Intelligence, New York, NY, USA, 2020: 12993-13000.
[40]
MaS, XuY. MPDIoU: a loss for efficient and accurate bounding box regression[J]. arXiv preprint arXiv:2023.
[41]
GeigerA, LenzP, StillerC, et al. Vision meets robotics: The KITTI dataset[J]. The International Journal of Robotics Research, 2013, 32(11): 1231-1237.
[42]
LiC, LiL, GengY, et al.YOLOv6 v3.0: a full-scale reloading[J]. arXiv preprint arXiv:2023.
WangC Y, YehI H, Mark LiaoH Y. YOLOv9: learning what you want to learn using programmable gradient information[C]∥European Conference on Computer Vision (ECCV), Milan, Italy, 2025: 1-21.
[45]
WangA, ChenH, LiuL, et al. YOLOv10: Real-time end-to-end object detection[J]. Advances in Neural Information Processing Systems (NeurIPS), 2024, 37: 107984-108011.
SzegedyC, VanhouckeV, IoffeS, et al. Rethinking the inception architecture for computer vision[C]∥IEEE Conference on Computer Vision and Pattern Recognition (CVPR),Las Vegas, Nevada, USA, 2016: 2818-2826.