Monocular 3D object detection encounters significant challenges due to depth ambiguity. Conventional 2D attention mechanisms have proven insufficient in mitigating this issue, and their high computational overhead further complicates deployment on vehicle-mounted mobile devices. To address these issues, this paper proposes a monocular 3D object detection algorithm based on 3D attention mechanism and multi-object bounding box module. Considering the depth ambiguity in 2D to 3D mapping, the 3D attention mechanism was first incorporated into the network design, which includes a depth information enhancement kernel and a position enhancement kernel with low computational complexity. Then, multi-object bounding box strategies employ pseudo-labels to alleviate the strict constraints of the original hard labels from depth labels perturbations. Hence, the precision of deep estimation is enhanced, improving the model’s 3D spatial perception and generalizability for 3D object detection tasks. Experiments on the nuScenes dataset demonstrate that this algorithm outperforms existing monocular 3D object detection methods. Finally, with the aid of TensorRT tools, the model is successfully deployed on mobile devices in the automotive environment. On the Jetson AGX Xavier and Jetson Orin NX (16 GB) embedded platforms, the inference time per frame is 67 ms and 89 ms, respectively, enabling real-time detection of 3D objects.
LINGC W, CHENH, XUD Y, et al .Self-supervised monocular depth estimation based on semantic assistance and depth temporal consistency constraints[J].Journal of Hunan University (Natural Sciences), 2024, 51(8): 1-12.(in Chinese)
[5]
WANGT, XINGEZ H U, PANGJ, et al. Probabilistic and geometric depth: detecting objects in perspective[C]//Conference on Robot Learning (CoRL). Auckland, New Zealand: PMLR, 2022: 1475-1485.
[6]
WANGG J, WUJ, TIANB, et al. CenterNet3D: an anchor free object detector for point cloud[J]. IEEE Transactions on Intelligent Transportation Systems, 2021, 23(8): 12953-12965.
[7]
LIUZ C, WUZ Z, TOTHR .SMOKE:single-stage monocular 3D object detection via keypoint estimation[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), June 14-19,2020. Seattle,WA,USA:IEEE,2020:4289-4298.
[8]
WANGT, ZHUX G, PANGJ M,et al .FCOS3D:fully convolutional one-stage monocular 3D object detection[C]//2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), October 11-17, 2021. Montreal,BC,Canada:IEEE,2021:913-922.
[9]
LIUX P, XUEN, WUT F. Learning auxiliary monocular contexts helps monocular 3D object detection[J]. Proceedings of the AAAI Conference on Artificial Intelligence,2022, 36(2): 1810-1818.
[10]
VASWANIA, SHAZEERN, PARMARN, et al. Attention is all you need[J]. Advances in Neural Information Processing Systems, 2017, 30: 5998-6008.
[11]
HUANGK, WUT, SUH,et al .MonoDTR:monocular 3D object detection with depth-aware transformer[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 18-24,2022. New Orleans,LA,USA:IEEE,2022:4002-4011.
[12]
ZHANGR R, QIUH, WANGT, et al. MonoDETR:depth-guided transformer for monocular 3D object detection[C]//2023 IEEE/CVF International Conference on Computer Vision (ICCV),October 1-6,2023. Paris,France:IEEE,2023: 9121-9132.
[13]
ZHOUY S, ZHUH Z, LIUQ, et al. MonoATT:online monocular 3D object detection with adaptive token transformer[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 17-24,2023.Vancouver, BC,Canada:IEEE, 2023: 17493-17503.
[14]
YANL F, YANP, XIONGS Z,et al .MonoCD:monocular 3D object detection with complementary depths[C]//2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 16-22,2024.Seattle,WA,USA:IEEE,2024:10248-10257.
[15]
LUY, MAX Z, YANGL,et al .Geometry uncertainty projection network for monocular 3D object detection[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV), October 10-17,2021. Montreal,QC,Canada:IEEE,2021:3091-3101.
[16]
SHIX P, YEQ, CHENX Z,et al .Geometry-based distance decomposition for monocular 3D object detection[C]//2021 IEEE/CVF International Conference on Computer Vision (ICCV),October 10-17,2021. Montreal,QC,Canada:IEEE,2021:15152-15161.
[17]
MOUSAVIANA, ANGUELOVD, FLYNNJ,et al .3D bounding box estimation using deep learning and geometry[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR),July 21-26,2017. Honolulu,HI,USA:IEEE,2017:5632-5640.
[18]
ZHANGY P, LUJ W, ZHOUJ .Objects are different:flexible monocular 3D object detection[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 20-25, 2021. Nashville, TN, USA: IEEE, 2021: 3288-3297.
[19]
CHENY J, TAIL, SUNK,et al .MonoPair:monocular 3D object detection using pairwise spatial relationships[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 13-19,2020. Seattle, WA, USA: IEEE, 2020:12090-12099.
[20]
XUH, GUOM T, NEDJAHN, et al .Vehicle and pedestrian detection algorithm based on lightweight YOLOv3-Promote and semi-precision acceleration[J].IEEE Transactions on Intelligent Transportation Systems, 2022, 23(10): 19760-19771.
[21]
DAIB, LIC, LINT,et al .Field robot environment sensing technology based on TensorRT[M]//Intelligent Robotics and Applications. Cham:Springer International Publishing,2021:370-377.
[22]
TANGY Z, QIANY. High-speed railway track components inspection framework based on YOLOv8 with high-performance model deployment[J]. High-speed Railway, 2024, 2(1): 42-50.
[23]
ZHANGJ F, SONGQ Z, YUJ. A transformer based complex-YOLOv4-trans for 3D point cloud object detection on embedded device[J]. Journal of Physics:Conference Series, 2022, 2404(1): 012026.
[24]
LIUZ J, TANGH T, AMINIA,et al .BEVFusion:multi-task multi-sensor fusion with unified bird’s-eye view representation[C]//2023 IEEE International Conference on Robotics and Automation (ICRA), May 29 - June 2, 2023. London,United Kingdom: IEEE, 2023: 2774-2781.
[25]
TAY Y, DEHGHANIM, BAHRID,et al .Efficient transformers:a survey[J].ACM Computing Surveys,2023,55(6):1-28.
[26]
CAIH, LIJ, HUM, et al. Efficientvit: multi-scale linear attention for high-resolution dense prediction [C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Paris, France: IEEE, 2023: 17302-17313.
LIUQ, LIW, YUS Y, et al. Monocular 3D object detection algorithm combining depth guidance and multi-scale channel attention mechanism[J]. Journal of Shandong University (Natural Science), 2025, 60(1): 63-73.(in Chinese)
[29]
HUANGC X, HET, RENH D,et al .OBMO:one bounding box multiple objects for monocular 3D object detection[J].IEEE Transactions on Image Processing,2023,32: 6570-6581.
[30]
NABATIR, QIH R .CenterFusion:center-based radar and camera fusion for 3D object detection[C]//2021 IEEE Winter Conference on Applications of Computer Vision (WACV),January 3-8,2021.Waikoloa,HI,USA:IEEE,2021:1526-1535.
[31]
YINT W, ZHOUX Y, KRAHENBUHLP .Center-based 3D object detection and tracking[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 20-25,2021. Nashville, TN, USA: IEEE, 2021: 11779-11788.
[32]
LANGA H, VORAS, CAESARH,et al .PointPillars:fast encoders for object detection from point clouds[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 15-20,2019. Long Beach,CA,USA:IEEE,2019:12689-12697.
[33]
SIMONELLIA, BULOS R, PORZIL,et al .Disentangling monocular 3D object detection[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV), October 27-November 2, 2019. Seoul, Korea: IEEE, 2019: 1991-1999.
[34]
WANGB L, ZHENGH W, ZHANGL,et al. BEVRefiner:improving 3D object detection in bird’s-eye-view via dual refinement[J]. IEEE Transactions on Intelligent Transportation Systems, 2024, 25(10): 15094-15105.
基金资助
国家自然科学基金资助项目(52072054)
National Natural ScienceFoundation of China(52072054)
重庆交通大学研究生科研创新项目(YYK202405)
Research and Innovation Program for Graduate Students in Chongqing Jiaotong University(YYK202405)