融合多尺度特征增强和代价体融合的三维重建方法

王云艳 ,  朱镇中 ,  熊超

湖北工业大学学报 ›› 2026, Vol. 41 ›› Issue (4) : 38 -43.

PDF (846KB)
湖北工业大学学报 ›› 2026, Vol. 41 ›› Issue (4) : 38 -43.

融合多尺度特征增强和代价体融合的三维重建方法

作者信息 +

Multi-view Stereo Network Reconstruction with Multi-scale Cross-Information Perception

Author information +
文章历史 +
PDF (866K)

摘要

为了提高多视图三维重建的精度,提出了一种基于多尺度交叉感知特征增强模块和多尺度代价体融合模块的重建方法.以CasMVSNet为基准网络,首先设计了一个交叉信息自注意力模块,利用纹理丰富的底层特征信息选择性地增强高层信息,将交叉信息自注意力模块嵌入U-Net网络中,采用高低层的语义特征融合方法,建立深层特征和浅层特征之间的依赖关系,提高特征提取的完整性,随后输出三种尺度不同大小的特征层;然后考虑到三种阶段的生成代价体相关性不足,设计一款代价体分层融合模块,将小尺度代价体中的信息分离并传递给下一层的代价体进行融合,由粗到精地进行深度估计,使重建精度和完整性都有大幅提高;最后考虑到粗阶段所生成的深度图中深度可能存在较大误差,会导致后续操作持续产生误差,因此在每一层后都添加了一个深度细化模块,使用参考图像作为指导来细化深度图,从而减少累积误差.由实验结果可知,该模型在DTU数据集上和基准网络CasMVSNet相比,在准确性误差和完整性误差上分别降低了6.89%和1.04%,相较于其他模型均有不同程度的提升,验证了该模型的有效性.

Abstract

To improve the accuracy of multi-view 3D reconstruction, a reconstruction method based on multi-scale cross perception feature enhancement module and multi-scale cost volume fusion module is proposed. The method takes CasMVSNet as the baseline network. Firstly, a cross information self attention module is designed to selectively enhance high-level information by incorporating it into the U-Net network and leveraging rich texture features from lower layers to improve the integrity of feature extraction. Three feature layers of different scales are then outputted to better capture multi-scale information. Secondly, considering the insufficient correlation of generated cost volumes in three stages, a cost volume hierarchical fusion module is designed. It separates and transfers information from small-scale cost volumes to the next layer for fusion, enabling depth estimation from coarse to fine, significantly improving reconstruction accuracy and completeness. Lastly, to address the potential large errors in depth maps generated by the coarse stage that could accumulate errors in subsequent operations, a depth refinement module is added after each layer. It refines the depth maps using reference images as guidance to reduce cumulative errors. Experimental results show that compared to the baseline network CasMVSNet and other models, the proposed method reduces accuracy error and completeness error by 6.89% and 1.04% respectively on the DTU dataset. This demonstrates the effectiveness of the approach, providing a strong reference for further research in the field of multi-view 3D reconstruction.

关键词

深度学习 / 交叉注意力 / 分层融合 / 深度细化 / 多视图立体匹配

Key words

deep learning / cross-attention / hierarchical fusion / depth refinement / multi-view stereo matching

引用本文

引用格式 ▾
王云艳,朱镇中,熊超. 融合多尺度特征增强和代价体融合的三维重建方法[J]. 湖北工业大学学报, 2026, 41(4): 38-43 DOI:

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

Lee H, Song S, Jo S. 3D reconstruction using a sparse laser scanner and a single camera for outdoor autonomous vehicle[C]// 2016 IEEE 19th International Conference on Intelligent Transportation Systems (ITSC). IEEE, 2016: 629-634.

[2]

Joseph S S, Aju D. A Comparative Survey on Three-Dimensional Reconstruction of Medical Modalities Based on Various Approaches[C]// Information Systems Design and Intelligent Applications: Proceedings of Fifth International Conference INDIA 2018 Volume 1. Springer Singapore, 2019: 223-233.

[3]

Sing K H, Xie W. Garden: A mixed reality experience combining virtual reality and 3D reconstruction[C]// Proceedings of the 2016 CHI Conference Extended Abstracts on Human Factors in Computing Systems. 2016: 180-183.

[4]

Nuchter A, Surmann H, Hertzberg J. Automatic model refinement for 3D reconstruction with mobile robots[C]// Fourth International Conference on 3-D Digital Imaging and Modeling, 2003. 3DIM 2003. Proceedings. IEEE, 2003: 394-401.

[5]

Peng K, Chen X, Zhou D, et al. 3D reconstruction based on SIFT and Harris feature points[C]// 2009 IEEE international conference on robotics and biomimetics (ROBIO). IEEE, 2009: 960-964.

[6]

Tafti A P, Baghaie A, Kirkpatrick A B, et al. A comparative study on the application of SIFT, SURF, BRIEF and ORB for 3D surface reconstruction of electron microscopy images[J]. Computer Methods in Biomechanics and Biomedical Engineering: Imaging & Visualization, 2018, 6(1): 17-30.

[7]

Ji M, Gall J, Zheng H, et al. SurfaceNet: An end-to-end 3d neural network for multiview stereopsis[C]// Proceedings of the IEEE international conference on computer vision. 2017: 2307-2315.

[8]

Yao Y, Luo Z, Li S, et al. Mvsnet: Depth inference for unstructured multi-view stereo[C]// Proceedings of the European conference on computer vision (ECCV). 2018: 767-783.

[9]

Yao Y, Luo Z, Li S, et al. Recurrent mvsnet for high-resolution multi-view stereo depth inference[C]// Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019: 5525-5534.

[10]

Yang J, Mao W, Alvarez J M, et al. Cost volume pyramid based depth inference for multi-view stereo[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020: 4877-4886.

[11]

Gu X, Fan Z, Zhu S, et al. Cascade cost volume for high-resolution multi-view stereo and stereo matching[C]// Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2020: 2495-2504.

[12]

Chen R, Han S, Xu J, et al. Point-based multi-view stereo network[C]// Proceedings of the IEEE/CVF international conference on computer vision. 2019: 1538-1547.

[13]

Yu Z, Gao S. Fast-mvsnet: Sparse-to-dense multi-view stereo with learned propagation and gauss-newton refinement[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020: 1949-1958.

[14]

Weilhartner R, Fraundorfer F. High res-mvsnet: A fast multi-view stereo network for dense 3d reconstruction from high-resolution images[J]. IEEE Access, 2021, 9: 11306-11315.

[15]

Zhang J, Li S, Luo Z, et al. Vis-mvsnet: Visibility-aware multi-view stereo network[J]. International Journal of Computer Vision, 2023, 131(1): 199-214.

[16]

Fu J, Liu J, Tian H, et al. Dual attention network for scene segmentation[C]// Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019: 3146-3154.

[17]

Wang G, Zhai Q, Liu H. Cross self-attention network for 3D point cloud[J]. Knowledge-Based Systems, 2022, 247: 108769.

[18]

Yang G, Manela J, Happold M, et al. Hierarchical deep stereo matching on high-resolution images[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019: 5515-5524.

AI Summary AI Mindmap
PDF (846KB)

0

访问

0

被引

详细

导航
相关文章

AI思维导图

/