摘要
针对3D目标检测模型中3D空间特征提取不充分的问题,提出一种改进的ICasA模型.首先,引入焦点稀疏卷积,对CasA模型的3D骨干网络进行改进,增强对非空体素特征位置之间信息流的提取,以获取更丰富的空间特征.其次,构建多尺度特征注意力融合模块,嵌入2D骨干网络,增强模型对多尺度特征的处理能力:采用不同步长卷积加深2D骨干网络结构,提升模型对全局特征的提取能力;改进注意力融合方式,在保留原始特征图的基础上实现对不同尺度特征不同位置的动态激活.在数据集KITTI上的实验结果表明,与CasA模型相比,ICasA模型对汽车、行人和骑行者的3D平均检测精度AP@R40分别提高0.28百分点、2.87百分点、1.03百分点,与现有先进模型相比,ICasA模型各类别目标检测效果更稳定,有助于提升3D目标检测的精度.
Abstract
Aiming at the problem of insufficient extraction of 3D spatial features in 3D object detection models, we proposed an improved ICasA model. Firstly, we introduced focal sparse convolution to improve the 3D backbone network of CasA model, strengthening the extraction of information flow among non-empty voxel feature locations to obtain richer spatial features. Secondly, we constructed a multi-scale feature attention fusion module and embedded it into the 2D backbone network to enhance the model’s ability to handle multi-scale features. We adopted convolutions with different strides to deepen the structure of the 2D backbone network, improve the model’s ability to extract global features. We improved the attention fusion method to dynamically activate features of different scales and locations while preserving the original feature maps. The experimental results on the KITTI dataset show that compared with the CasA model, the ICasA model achieves 0.28 percentage points, 2.87 percentage points, and 1.03 percentage points increases in 3D average detection accuracy (AP@R40) for cars, pedestrians, and cyclists, respectively. Compared with existing advanced models, the ICasA model has more stable detection effect for all categories of objects, which helps to enhance the accuracy of 3D object detection.
关键词
Key words
[Author(id=1291050201008915179, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, orderNo=0, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=hqhe@cauc.edu.cn, emailSecond=null, emailThird=null, correspondingAuthor=0, authorType=1, ext={EN=AuthorExt(id=1291050201067635437, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, authorId=1291050201008915179, language=EN, stringName=Huaiqing He, firstName=Huaiqing, middleName=null, lastName=He, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=College of Computer and Artificial Intelligence, Civil Aviation University of China, Tianjin 300300, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1291050201117967086, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, authorId=1291050201008915179, language=CN, stringName=贺怀清, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=中国民航大学 计算机与人工智能学院, 天津 300300, bio={"content":"贺怀清(1969—),女,汉族,博士,教授,从事图形图像和可视分析的研究, E-mail: hqhe@cauc.edu.cn.
"}, bioImg=null, bioContent=贺怀清(1969—),女,汉族,博士,教授,从事图形图像和可视分析的研究, E-mail: hqhe@cauc.edu.cn.
, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1291050200941806311, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, xref=null, ext=[AuthorCompanyExt(id=1291050200954389224, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, companyId=1291050200941806311, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=College of Computer and Artificial Intelligence, Civil Aviation University of China, Tianjin 300300, China), AuthorCompanyExt(id=1291050200966972137, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, companyId=1291050200941806311, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=中国民航大学 计算机与人工智能学院, 天津 300300)])]), Author(id=1291050201159910129, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, orderNo=1, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=zhaiyujia_wish@163.com, emailSecond=null, emailThird=null, correspondingAuthor=1, authorType=1, ext={EN=AuthorExt(id=1291050201214436084, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, authorId=1291050201159910129, language=EN, stringName=Yujia Zhai, firstName=Yujia, middleName=null, lastName=Zhai, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=College of Computer and Artificial Intelligence, Civil Aviation University of China, Tianjin 300300, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1291050201310905077, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, authorId=1291050201159910129, language=CN, stringName=翟羽佳, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=中国民航大学 计算机与人工智能学院, 天津 300300, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1291050200941806311, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, xref=null, ext=[AuthorCompanyExt(id=1291050200954389224, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, companyId=1291050200941806311, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=College of Computer and Artificial Intelligence, Civil Aviation University of China, Tianjin 300300, China), AuthorCompanyExt(id=1291050200966972137, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, companyId=1291050200941806311, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=中国民航大学 计算机与人工智能学院, 天津 300300)])]), Author(id=1291050201357042423, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, orderNo=2, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=null, emailSecond=null, emailThird=null, correspondingAuthor=0, authorType=1, ext={EN=AuthorExt(id=1291050201415762681, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, authorId=1291050201357042423, language=EN, stringName=Haohan Liu, firstName=Haohan, middleName=null, lastName=Liu, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=College of Computer and Artificial Intelligence, Civil Aviation University of China, Tianjin 300300, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1291050201461900028, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, authorId=1291050201357042423, language=CN, stringName=刘浩翰, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=中国民航大学 计算机与人工智能学院, 天津 300300, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1291050200941806311, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, xref=null, ext=[AuthorCompanyExt(id=1291050200954389224, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, companyId=1291050200941806311, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=College of Computer and Artificial Intelligence, Civil Aviation University of China, Tianjin 300300, China), AuthorCompanyExt(id=1291050200966972137, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, companyId=1291050200941806311, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=中国民航大学 计算机与人工智能学院, 天津 300300)])]), Author(id=1291050201524814594, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, orderNo=3, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=null, emailSecond=null, emailThird=null, correspondingAuthor=0, authorType=1, ext={EN=AuthorExt(id=1291050201587729158, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, authorId=1291050201524814594, language=EN, stringName=Kanghua Hui, firstName=Kanghua, middleName=null, lastName=Hui, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=College of Computer and Artificial Intelligence, Civil Aviation University of China, Tianjin 300300, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1291050201633866504, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, authorId=1291050201524814594, language=CN, stringName=惠康华, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=null, address=中国民航大学 计算机与人工智能学院, 天津 300300, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1291050200941806311, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, xref=null, ext=[AuthorCompanyExt(id=1291050200954389224, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, companyId=1291050200941806311, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=College of Computer and Artificial Intelligence, Civil Aviation University of China, Tianjin 300300, China), AuthorCompanyExt(id=1291050200966972137, tenantId=1045748351789510663, journalId=1155139928303341642, articleId=1291050200023253729, companyId=1291050200941806311, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=中国民航大学 计算机与人工智能学院, 天津 300300)])])]
贺怀清,翟羽佳,刘浩翰,惠康华.
基于特征增强的3D目标检测模型ICasA[J].
吉林大学学报(理学版), 2026, 64(4): 823-834 DOI:10.13413/j.cnki.jdxblxb.2025157
| [1] |
杨笑天, 谭金林, 鱼昕, 等. 基于GAM-YOLOv8的遥感图像舰船目标跟踪[J]. 吉林大学学报(地球科学版), 2025, 55(1):328-339.
|
| [2] |
(Yang X T, Tan J L, Yu X, et al. Ship Target Tracking Based on GAM-YOLOv8 Remote Sensing Images[J]. Journal of Jilin University(Earth Science Edition), 2025, 55(1):328-339.)
|
| [3] |
Lin X, Peng J L, Gan Z Y, et al. YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-Time Detection[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2026: 18440-18449.
|
| [4] |
唐传茵, 吴龙杰, 於涛, 等. 不利天气条件下目标检测的传感器融合方法[J]. 东北大学学报(自然科学版), 2026, 47(2):58-65.
|
| [5] |
(Tang C Y, Wu L J, Yu T, et al. Target Detection Sensor Fusion Method under Adverse Weather Conditions[J]. Journal of Northeastern University(Natural Science), 2026, 47(2):58-65.)
|
| [6] |
Chen X Z, Ma H M, Wan J, et al. Multi-view 3D Object Detection Network for Autonomous Driving[C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2017: 1907-1915.
|
| [7] |
Ku J, Mozifian M, Lee J, et al. Joint 3D Proposal Generation and Object Detection from View Aggregation[C]// 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems. Piscataway, NJ: IEEE, 2018: 1-8.
|
| [8] |
Wu X P, Peng L, Yang H H, et al. Sparse Fuse Dense: Towards High Quality 3D Detection with Depth Completion[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2022: 5418-5427.
|
| [9] |
Li X, Ma T, Hou Y N, et al. LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross-Modal Fusion[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2023: 17524-17534.
|
| [10] |
田枫, 宗内丽, 刘芳, 等. 多模态融合的三维目标检测方法研究[J]. 计算机工程与应用, 2024, 60(13):113-123.
|
| [11] |
(Tian F, Zong N L, Liu F, et al. Research on 3D Object Detection Method Based on Multi-modal Fusion[J]. Computer Engineering and Applications, 2024, 60(13): 113-123.)
|
| [12] |
Chen Y K, Li Y W, Zhang X Y, et al. Focal Sparse Convolutional Networks for 3D Object Detection[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2022: 5428-5437.
|
| [13] |
Graham B, Engelcke M, Van Der Maaten L. 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks[C]// Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2018: 9224-9232.
|
| [14] |
Lu T, Ding X, Liu H S, et al. LinK: Linear Kernel for LiDAR-Based 3D Perception[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2023: 1105-1115.
|
| [15] |
Wu H, Deng J H, Wen C L, et al. CasA: A Cascade Attention Network for 3-D Object Detection from LiDAR Point Clouds[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 1-11.
|
| [16] |
Chen Y D, Cai G R, Xia Q M, et al. EPDet: Enhancing Point Clouds Features with Effective Representation for 3D Object Detection[J]. International Journal of Applied Earth Observation and Geoinformation, 2024, 127: 103688-1-103688-12.
|
| [17] |
Zheng W L, Tang W, Chen S J, et al. CIA-SSD: Confident IoU-Aware Single-Stage Object Detector from Point Cloud[C]// Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto: AAAI Press, 2021: 3555-3562.
|
| [18] |
Geiger A, Lenz P, Urtasun R. Are We Ready for Autonomous Driving? The KITTI Vision Benchmark Suite[C]// 2012 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2012: 3354-3361.
|
| [19] |
Shi S S, Guo C X, Jiang L, et al. PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object Detection[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2020: 10529-10538.
|
| [20] |
Deng J J, Shi S S, Li P W, et al. VoxelR-CNN: Towards High Performance Voxel-Based 3D Object Detection[C]// Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto: AAAI Press, 2021: 1201-1209.
|
| [21] |
Sheng H L, Cai S J, Liu Y, et al. Improving 3D Object Detection with Channel-Wise Transformer[C]// Proceedings of the IEEE/CVF International Conference on Computer Vision. Piscataway, NJ: IEEE, 2021: 2743-2752.
|
| [22] |
Zheng W, Tang W L, Jiang L, et al. SE-SSD: Self-ensembling Single-Stage Object Detector from Point Cloud[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2021: 14494-14503.
|
| [23] |
Xu Q G, Zhong Y Q, Neumann U. Behind the Curtain Learning Occluded Shapes for 3D Object Detection[C]// Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto: AAAI Press, 2022: 2893-2901.
|
| [24] |
Zhang Y N, Chen J X, Huang D. CAT-Det: Contrastively Augmented Transformer for Multi-modal 3D Object Detection[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2022: 908-917.
|
| [25] |
Zhou C, Zhang Y N, Chen J Y, et al. OcTr: Octree-Based Transformer for 3D Object Detection[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2023: 5166-5175.
|
| [26] |
Mahmoud A, Hu J S K, Waslander S L. Dense Voxel Fusion for 3D Object Detection[C]// Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. Piscataway, NJ: IEEE, 2023: 663-672.
|
| [27] |
Zhang T Y, Liang Z G, Yang Y Z, et al. Contrastive Late Fusion for 3D Object Detection[J]. IEEE Transactions on Intelligent Vehicles, 2025, 10(5): 3442-3457.
|
| [28] |
Li Z, Yang Z J, Gao Y L, et al. 6DoF-3D: Efficient and Accurate 3D Object Detection Using Six Degrees-of-Freedom for Autonomous Driving[J]. Expert Systems with Applications, 2024, 238: 122319-1-122319-14.
|
基金资助
国家自然科学基金(U1333110)
国家重点研发计划项目(2020YFB1600101)
天津市教委科研项目(2020KJ024)