To improve the depth estimation accuracy in complex and changeable scenes, a new monocular depth estimation network based on U-shaped encoder-decoder is proposed. The Swin Transformer architecture is adopted as the core of the encoder to realize fine-grained feature extraction of input data at multiple levels and scales. Multi-scale local features are extracted by using layer-by-layer dilated convolution. The local and global features are interacted through the feature interaction module to achieve a more comprehensive understanding of complex scenes. A symmetric transformer decoder is adopted and combined with an image patch expansion layer to reshape the feature map of adjacent dimensions into a feature map with higher resolution. Eventually, pixel-level depth prediction is output. Quantitative experiments are conducted on the NYU Depth v2 dataset and the KITTI dataset. The research results show that this network has high efficiency and practicability in complex and changeable scenes. The research conclusion breaks through the limitations of traditional methods in complex and changeable scenes and provides new perspectives and methodologies for the theoretical research of depth estimation.
单目深度估计是一种通过单个摄像头拍摄的图像来推断场景中各物体深度信息的技术。从图像中估计深度信息是计算机视觉的一项基础且重要的功能,可广泛应用于同步定位与建图(simultaneous localization and mapping,SLAM)[1]、三维重建[2]、目标检测[3]和语义分割[4] 等领域。
NYU Depth v2数据集的定性结果见图6。由图6可知,在捕捉和再现较小空间物体的锐利边缘与细微纹理时,本文方法的精确度较高。对于灯杆类细长且纹理单一的物体,其他方法因分辨率或光照条件限制,出现深度估计模糊或失真问题,而本文方法注重轮廓细节还原,能清晰地勾勒灯杆轮廓、准确捕捉桌子边缘结构,并精细还原桌面上的细微纹理和阴影变化。针对椅子脚等带黑色弱纹理待估计对象,本文方法通过良好学习能力推断补全椅子脚的深度信息,展现出强大的深度估计能力。
LIUDong, YUTao, CONGMing,et al.Visual SLAM method for dynamic environment based on deep learning image features[J].Journal of Huazhong University of Science and Technology (Natural Science Edition),2024,52(6):156-163.
GUOBaoyun, YAOYukai, LICailin,et al.Application of improved 3D-BoNet to segmentation and 3D reconstruction of point cloud instances[J].Bulletin of Surveying and Mapping,2024(6):30-35.
ZHANGZhijia, FANYingying, SHAOYiming,et al.Detection of multi-type traffic signs based on improved YOLO v3 model[J].Journal of Shenyang University of Technology,2023,45(1):66-70.
LIUJingwei, ZHOUYan.Local feature aggregation networks for 3D semantic segmentation[J].Computing Technology and Automation,2024, 43(2):170-176.
[9]
LAINAI, RUPPRECHTC, BELAGIANNISV,et al.Deeper depth prediction with fully convolutional residual networks[C]//2016 Fourth International Conference on 3D Vision (3DV).October 25-28,2016,Stanford,CA,USA.IEEE,2016:239-248.
[10]
YINZ C, SHIJ P.GeoNet: unsupervised learning of dense depth, optical flow and camera pose[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. June 18-23,2018,Salt Lake City,UT,USA.IEEE,2018:1983-1992.
YINYameng, ZHOUJiaqi, WANGZhihui.Research on monocular depth estimation based on diverse branch block[J].Computer & Digital Engineering,2023,51(12):2966-2970.
[13]
SUIX, GAOS, XUA G,et al.Lightweight monocular depth estimation using a fusion-improved transformer[J].Scientific Reports,2024,14(1): 22472.
ZHENGYuhang, CAOChuqing.Continuous frame depth estimation based on multi-scale feature mixed attention mechanism[J].Journal of Chongqing Technology and Business University (Natural Science Edition),2024,41(4):104-111.
LITao, HUTing, WUDandan.Monocular depth estimation combining pyramid structure and attention mechanism[J].Journal of Graphics, 2024,45(3):454-463.
[19]
EIGEND, PUHRSCHC, FERGUSR.Depth map prediction from a single image using a multi-scale deep network[C]//Proceedings of the 28th International Conference on Neural Information Processing Systems.Cambridge,Massachusetts:MIT Press,2014:2366-2374.
[20]
FUH, GONGM M, WANGC H,et al.Deep ordinal regression network for monocular depth estimation[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition.June 18-23,2018,Salt Lake City,UT,USA.IEEE,2018:2002-2011.
[21]
RAMAMONJISOAM, LEPETITV.SharpNet:fast and accurate recovery of occluding contours in monocular depth estimation[C]//2019 IEEE/CVF International Conference on Computer Vision Workshop. October 27-28,2019,Seoul,Korea (South).IEEE,2019:2109-2118.
[22]
PATILV, SAKARIDISC, LINIGERA,et al.P3Depth:monocular depth estimation with a piecewise planarity prior[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition.June 18-24, 2022,New Orleans,LA,USA.IEEE,2022:1600-1611.
[23]
FAROOQ BHATS, ALHASHIMI, WONKAP.AdaBins:depth estimation using adaptive bins[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition. June 20-25,2021,Nashville, TN,USA.IEEE,2021:4008-4017.
[24]
BHATS F, ALHASHIMI, WONKAP.LocalBins:improving depth estimation by Learning local distributions[C]//Computer Vision-ECCV2022.Cham:Springer Nature Switzerland,2022:480-496.
[25]
RANFTLR, BOCHKOVSKIYA, KOLTUNV.Vision transformers for dense prediction[C]//2021 IEEE/CVF International Conference on Computer Vision.October 10-17,2021,Montreal,QC,Canada.IEEE, 2021:12159-12168.
[26]
YUANW H, GUX D, DAIZ Z,et al.Neural window fully-connected CRFs for monocular depth estimation[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition.June 18-24,2022,New Orleans,LA,USA.IEEE,2022:3906-3915.
[27]
LONGX X, LINC, LIUL J,et al.Adaptive surface normal constraint for depth estimation[C]//2021 IEEE/CVF International Conference on Computer Vision.October 10-17,2021,Montreal,QC,Canada.IEEE,2021:12829-12838.
[28]
LEES, LEEJ, KIMB,et al.Patch-wise attention network for monocular depth estimation[J].Proceedings of the AAAI Conference on Artificial Intelligence,2021,35(3):1873-1881.
[29]
LIZ Y, CHENZ H, LIUX M,et al.DepthFormer:exploiting long-range correlation and local information for accurate monocular depth estimation[J].Machine Intelligence Research,2023,20(6):837-854.
[30]
AGARWALA, ARORAC.Attention attention everywhere:monocular depth prediction with skip attention[C]//2023 IEEE/CVF Winter Conference on Applications of Computer Vision.January 2-7,2023, Waikoloa,HI,USA. IEEE,2023:5850-5859.