基于多层级交互式特征融合的三维目标检测算法

高凯 ,  王晟宇 ,  付强 ,  才华 ,  张晨洁 ,  王伟刚

吉林大学学报(理学版) ›› 2026, Vol. 64 ›› Issue (3) : 591 -602.

PDF (4462KB)
吉林大学学报(理学版) ›› 2026, Vol. 64 ›› Issue (3) : 591 -602. DOI: 10.13413/j.cnki.jdxblxb.2025021
计算机科学

基于多层级交互式特征融合的三维目标检测算法

作者信息 +

Three-Dimensional Object Detection Algorithm Based on Multi-level Interactive Feature Fusion

Author information +
文章历史 +
PDF (4568K)

摘要

针对自动驾驶场景中三维目标检测存在的小目标识别困难、远距离点云稀疏以及多模态特征融合不足等问题,在多模态三维目标检测框架基础上提出一种改进算法.该算法通过构建类别与质心感知的前景点采样策略,增强前景信息保留能力并抑制背景噪声干扰;通过引入动态卷积图像特征提取机制,提高图像特征表达质量;通过设计多阶段交互式特征注意力融合模块,提升点云与图像特征的深层协同建模能力.实验结果表明,该算法在公开数据集上对汽车、行人和骑行者三类目标的平均检测精度分别达83.49%,46.98%和68.28%,整体性能优于当前主流方法.该方法能有效提升复杂交通场景下三维目标检测的准确性和鲁棒性,对推动自动驾驶环境感知技术的发展有一定参考价值.

Abstract

Aiming at the problems of small-object recognition difficulty, sparse point clouds at long distances, and insufficient multimodal feature fusion in three-dimensional object detection for autonomous driving scenarios, we proposed an improved algorithm based on a multimodal three-dimensional object detection framework. The algorithm enhanced the ability to preserve foreground information and suppress background noise interference by constructing a class-and centroid-aware foreground point sampling strategy. By introducing a dynamic convolutional image feature extraction mechanism, the quality of image feature representation was improved. By designing a multi-stage interactive feature attention fusion module, the deep collaborative modeling ability between point cloud features and image features was improved. Experimental results on a public dataset show that the proposed method achieves average detection accuracies of 83.49%, 46.98% and 68.28% for three types of objects: cars, pedestrians and cyclists, respectively, and outperforms current mainstream methods in overall performance. The proposed method can effectively improve the accuracy and robustness of three-dimensional object detection in complex traffic scenarios and has certain reference value for promoting the development of autonomous driving environment perception technology.

关键词

计算机视觉 / 关键点采样 / 动态卷积 / 三维目标检测 / 多模态融合

Key words

computer vision / keypoint sampling / dynamic convolution / 3D object detection / multimodal fusion

引用本文

引用格式 ▾
高凯,王晟宇,付强,才华,张晨洁,王伟刚. 基于多层级交互式特征融合的三维目标检测算法[J]. 吉林大学学报(理学版), 2026, 64(3): 591-602 DOI:10.13413/j.cnki.jdxblxb.2025021

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

STEIN G P, MANO O, SHASHUA A. Vision-Based ACC with a Single Camera: Bounds on Range and Range Rate Accuracy[C]// IEEE IV 2003 Intelligent Vehicles Symposium. Piscataway, NJ: IEEE, 2003: 120-125.

[2]

QIN Z Y, WANG J L, LU Y. Monogrnet: A Geometric Reasoning Network for Monocular 3D Object Localization[C]// Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto: AAAI Press, 2019: 8851-8858.

[3]

CHEN X Z, KUNDU K, ZHANG Z Y, et al. Monocular 3D Object Detection for Autonomous Driving[C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2016: 2147-2156.

[4]

HE T, SOATTO S. Mono3D++: Monocular 3D Vehicle Detection with Two-Scale 3D Hypotheses and Task Priors[C]// Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto: AAAI Press, 2019: 8409-8416.

[5]

刘长吉. 以二维图像驱动的三维目标检测[D]. 长春: 中国科学院大学(中国科学院长春光学精密机械与物理研究所), 2021.

[6]

(LIU C J. 3D Object Detection Driven by 2D Images[D]. Changchun: University of Chinese Academy of Sciences (Changchun Institute of Optics, Fine Mechanics and Physics, Chinese Academy of Sciences), 2021.)

[7]

CHEN X Z, KUNDU K, ZHU Y K, et al. 3D Object Proposals Using Stereo Imagery for Accurate Object Class Detection[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 40(5): 1259-1272.

[8]

YANG Z T, SUN Y N, LIU S, et al. STD: Sparse-to-Dense 3D Object Detector for Point Cloud[C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Piscataway, NJ: IEEE, 2019: 1951-1960.

[9]

才华, 寇婷婷, 杨依宁, . 基于轨迹优化的三维车辆多目标跟踪[J]. 吉林大学学报(工学版), 2024, 54(8):2338-2347.

[10]

(CAI H, KOU T T, YANG Y N, et al. 3D Vehicle Multi-object Tracking Based on Trajectory Optimization[J]. Journal of Jilin University (Engineering and Technology Edition), 2024, 54(8): 2338-2347.)

[11]

ZHOU Y, TUZEL O. Voxelnet: End-to-End Learning for Point Cloud Based 3D Object Detection[C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2018: 4490-4499.

[12]

SHI S S, WANG X G, LI H S. Pointrcnn: 3D Object Proposal Generation and Detection from Point Cloud[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2019: 770-779.

[13]

SHI S S, GUO C X, JIANG L, et al. PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object Detection[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2020: 10529-10538.

[14]

才华, 郑延阳, 付强, . 基于多尺度候选融合与优化的三维目标检测算法[J]. 吉林大学学报(工学版), 2025, 55(2):709-721.

[15]

(CAI H, ZHENG Y Y, FU Q, et al. 3D Object Detection Algorithm Based on Multi-scale Candidate Fusion and Optimization[J]. Journal of Jilin University (Engineering and Technology Edition), 2025, 55(2): 709-721.)

[16]

YANG Z T, SUN Y N, LIU S, et al. 3DSSD: Point-Based 3D Single Stage Object Detector[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2020: 11040-11048.

[17]

QI C R, SU H, MO K, et al. Pointnet: Deep Learning on Point Sets for 3D Classification and Segmentation[C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2017: 652-660.

[18]

QI C R, YI L, SU H, et al. Pointnet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space[J]. Advances in Neural Information Processing Systems, 2017, 30: 1-10.

[19]

QI C R, LIU W, WU C X, et al. Frustum Pointnets for 3D Object Detection from RGB-D Data[C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2018: 918-927.

[20]

CHEN X Z, MA H M, WAN J, et al. Multi-view 3D Object Detection Network for Autonomous Driving[C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2017: 1907-1915.

[21]

KU J, MOZIFIAN M, LEE J, et al. Joint 3D Proposal Generation and Object Detection from View Aggregation[C]// 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). Piscataway, NJ: IEEE, 2018: 1-8.

[22]

VORA S, LANG A H, HELOU B, et al. Point painting: Sequential Fusion for 3D Object Detection[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2020: 4604-4612.

[23]

YOO J H, KIM Y, KIM J, et al. 3D-CVF: Generating Joint Camera and Lidar Features Using Cross-View Spatial Feature Fusion for 3D Object Detection[C]// 16th European Conference on Computer Vision. Berlin: Springer International Publishing, 2020: 720-736.

[24]

PANG S, MORRIS D, RADHA H. CLOCs: Camera-LiDAR Object Candidates Fusion for 3D Object Detection[C]// 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). Piscataway, NJ: IEEE, 2020: 10386-10393.

[25]

LIU Z, HUANG T T, LI B L, et al. Epnet++: Cascade Bi-directional Fusion for Multi-modal 3D Object Detection[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 45(7): 8324-8341.

[26]

HUANG T T, LIU Z, CHEN X W, et al. Epnet: Enhancing Point Features with Image Semantics for 3D Object Detection[C]// 16th European Conference on Computer Vision. Berlin: Springer International Publishing, 2020: 35-52.

[27]

LIN T Y, GOYAL P, GIRSHICK R, et al. Focal Loss for Dense Object Detection[C]// Proceedings of the IEEE International Conference on Computer Vision. Piscataway, NJ: IEEE, 2017: 2980-2988.

[28]

GEIGER A, LENZ P, URTASUN R. Are We Ready for Autonomous Driving? The KITTI Vision Benchmark Suite[C]// 2012 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2012: 3354-3361.

基金资助

国家自然科学基金联合基金(U2341226)

吉林省科技厅项目(20260102267JC)

吉林省科技厅项目(20240302089GX)

吉林省医疗卫生人才专项基金(JLSWSRCZX2023-70)

2024年空间智能控制技术全国重点实验室开放基金(2024-CXPT-GF-JJ-012-12)

AI Summary AI Mindmap
PDF (4462KB)

101

访问

0

被引

详细

导航
相关文章

AI思维导图

/