Adopting single modal to detect pedestrians has a high missed detection rate under the circumstance of poor illumination and partial occlusion,while dual-modal has low real-time detection speed. To address this issue, this study proposed a lightweight method using dual-modal to detect pedestrians based on an improved deep learning model. Firstly, a Semantic-Aware Real-Time Image Fusion Network (SeAFusion)based on semantic perception was employed to fuse infrared and visible-light images at the feature level. Secondly, a lightweight network, GhostNet, was integrated into the backbone network of the YOLOv5 baseline model. Thirdly, following the Spatial Pyramid Pooling Fast (SPPF) module in the backbone network, a Global Attention Mechanism (GAM) was added to adaptively acquire global multi-dimensional features and further reduce information diffusion. To improve the convergence speed and accuracy of model solving,the Complete Intersection Over Union (CIoU) loss function was replaced by the Efficient Intersection Over Union (EIoU). Finally, this study employed the open dataset LLVIP for experimental validation. The results show that compared with the unimproved YOLOv5 model, the proposed method can increase the image processing rate by 4 frames per second and achieve a mean average precision of 94.7% for pedestrian detection with a 32.1% reduction in model weights.
LiuJ J, ZhangS T, WangS, et al. Multispectral deep neural networks for pedestrian detection[C]∥Proceedings of the British Machine Vision Conference, York, UK: BMVC, 2016.
GuoXiao-han, PengLi-qun, MaDing-hui. A method of identifying collision risk of container trucks in port terminal areas under an integrated connected vehicle BSM and roadside video surveillance data[J]. Journal of Transport Information and Safety, 2023, 41(1): 1-12.
ZhaoBin, WangChun-ping, FuQiang, et al. Multi-scale infraredpedestrian detection based on deep attention mechanism[J]. Acta Optica Sinica, 2020, 40(5): 47-58.
[6]
ZhuangY F, PuZ Y, HuJ, et al. Illumination and temperature-aware multispectral networks for edge-computing enabled pedestrian detection[J]. IEEE Transactions on Network Science and Engineering, 2022, 9(3): 1282-1295.
YouFeng, LiangJian-zhong, CaoShui-jin, et al. Dense pedestrian crowd trajectory extraction and motion semantic perceptionbased on multi-object tracking[J]. Journal of Transportation Systems Engineering and Information Technology, 2021, 21 (6): 42-54, 95.
[9]
ViolaP, JonesM J. Robust real-time face detection[J]. International Journal of Computer Vision, 2001, 57(2): 137-154.
[10]
DalalN, TriggsB. Histograms of oriented gradients for human detection[C]∥IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Diego, CA, USA: IEEE, 2005.
ChengXu, SongChen, ShiJin-gang, et al. A survey of generic object detection methods based on deep learning[J]. Acta Electronica Sinica, 2021, 49(7): 1428-1438.
[13]
RenS Q, HeK M, GirshickR, et al. Faster R-CNN: towards real-time object detection with region proposal networks[J].IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(6): 1137-1149.
[14]
HeK M, GkioxariG, DollárP, et al. Mask R-CNN[C]∥IEEE International Conference on Computer Vision (ICCV), Venice, Italy: IEEE, 2017.
[15]
LiuW, AnguelovD, ErhanD, et al. SSD: single shot multibox detector[C]∥European Conference on Computer Vision. Amsterdam, The Netherlands: ECCV, 2016.
[16]
AdarshP, RathiP, KumarM. YOLO v3-Tiny: object detection and recognition using one stage improved model[C]∥2020 6th International Conference on Advanced Computing and Communication Systems (ICACCS). Coimbatore, India: IEEE, 2020.
LuoYan, ZhangChong-yang, TianYong-hong, et al. An overview of deep learning based pedestrian detection algorithms[J]. Journal of Image and Graphics, 2022, 27(7): 2094-2111.
[19]
LiC Y, SongD, TongR F, et al. Illumination-aware faster R-CNN for robust multispectral pedestrian detection[J]. Pattern Recognition, 2019, 85:161-171.
[20]
GuanD Y, CaoY P, LiangJ, et al. Fusion of multispectral data through illumination-aware deep neural networks for pedestrian detection[J]. Information Fusion, 2019, 50:148-157.
[21]
XueY J, JuZ Y, LiY M, et al. MAF-YOLO: Multi-modal attention fusion based YOLO for pedestrian detection[J]. Infrared Physics & Technology, 2021, 118: 103906.
[22]
LiC, WangY D, LiuX M. A multi-pedestrian tracking algorithm for dense scenes based on an attention mechanism and dual data association[J]. Applied Sciences, 2022, 12(19): 9597-9607.
SunWen-cai, HuXu-ge, YangZhi-fa, et al. Optimization of infrared-visible road target detection by fusing GPNet and image multiscale features[J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(10): 2799-2806.
WangHong-zhi, SongMing-xuan, ChengChao, et al. Road object detection method based on improved YOLOv5 algorithm[J]. Journal of Jilin University (Engineering and Technology Edition), 2024, 54(9): 2658-2667.
[30]
TangL F, YuanJ T, MaJ Y. Image fusion in the loop of high-level vision tasks: a semantic-aware real-time infrared and visible image fusion network[J]. Information Fusion, 2022, 82: 28-42.
[31]
HanK, WangY H, TianQ, et al. Ghostnet: more features from cheap operations[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle, WA, USA: IEEE, 2020.
[32]
LiuY C, ShaoZ R, HoffmannN. Global attention mechanism: retain information to enhance channel-spatial interactions[EB/OL]. 2024-11-10.
[33]
JiaX Y, ZhuC, LiM Z, et al. LLVIP: a visible-infrared paired dataset for low-light vision[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal, QC, Canada: IEEE, 2021.