This paper proposes a novel depth and pose estimation framework, leveraging the wavelet transform and the self-supervised structure from the motion paradigm. The approach involves embedding 2D discrete wavelet transform into neural networks and implementing gradient propagation. Traditional convolutional neural networks (CNN) face a challenge during the down-sampling stage, as structural information is lost, and becomes irrecoverable in subsequent phases. This loss of information impacts the performance of depth estimation tasks, where complete structural information is crucial. This paper uses a 2D discrete wavelet transform layer to replace the down-sampling process of traditional neural networks, which can better preserve the structural details and avoid the accumulation of noise. In the up-sampling stage of the decoder, the inverse wavelet transform layer is used to replace the conventional interpolation method, which can effectively restore detailed information and promote the accuracy of the depth map. In addition, the proposed method has noise robustness compared to traditional neural networks. Experiments on the KITTI dataset demonstrate that the proposed algorithm performs excellently in the self-supervised depth and pose estimation tasks.
HARTLEYR, ZISSERMANA. Multiple View Geometry in Computer Vision[M]. Cambridge: Cambridge University Press, 2000.
[2]
EIGEND, PUHRSCHC, FERGUSR. Depth map prediction from a single image using a multi-scale deep network[EB/OL]. 2014: arXiv: 1406.2283.
[3]
MAYERN, ILG E, FISCHERP, et al. What makes good synthetic training data for learning disparity and optical flow estimation [J]. International Journal of Computer Vision, 2018, 126(9): 942-960. DOI:10.1007/s11263-018-1082-6 .
[4]
ZHOUT H, BROWNM, SNAVELYN, et al. Unsupervised learning of depth and ego-motion from video[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2017: 6612-6619. DOI:10.1109/CVPR.2017.700 .
[5]
GARGR, VIJAY KUMARB G, CARNEIROG, et al. Unsupervised CNN for single view depth estimation: Geometry to the rescue[C]// 2016 European Conference on Computer Vision. Cham: Springer International Publishing, 2016:740-756. DOI:10.1007/978-3-319-46484-8_45 .
[6]
YINZ C, SHIJ P. GeoNet: Unsupervised learning of dense depth, optical flow and camera pose[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 1983-1992. DOI:10.1109/CVPR.2018.00212 .
[7]
GODARDC, AODHAO M, FIRMANM, et al. Digging into self-supervised monocular depth estimation[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE Press, 2019: 3827-3837. DOI:10.1109/ICCV.2019.00393 .
YEX Y, HEY L, RUS N. Unsupervised monocular depth estimation and visual odometry based on generative adversarial network and self-attention mechanism[J]. Robot, 2021, 43(2): 203-213. DOI:10.13973/j.cnki.robot.200084(Ch ).
[12]
BIANJ W, LIZ C, WANGN Y, et al. Unsupervised scale-consistent depth and ego-motion learning from monocular video[EB/OL]. 2019: arXiv: 1908.10553.
[13]
DUANY P, LIUF, JIAOL C, et al. SAR image segmentation based on convolutional-wavelet neural network and Markov random field[J]. Pattern Recognition, 2017, 64: 255-267. DOI:10.1016/j.patcog.2016.11.015 .
[14]
LIUP J, ZHANGH Z, ZHANGK, et al. Multi-level wavelet-CNN for image restoration[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). New York: IEEE Press, 2018: 886-88609. DOI:10.1109/CVPRW.2018.00121 .
[15]
HEK M, ZHANGX Y, RENS Q, et al. Deep residual learning for image recognition[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2016: 770-778. DOI:10.1109/CVPR.2016.90 .
[16]
WANGZ, BOVIKA C, SHEIKHH R, et al. Image quality assessment: From error visibility to structural similarity[J]. IEEE Transactions on Image Processing, 2004, 13(4): 600-612. DOI:10.1109/TIP.2003.819861 .
[17]
GEIGERA, LENZP, STILLERC, et al. Vision meets robotics: The KITTI dataset[J]. The International Journal of Robotics Research, 2013, 32(11): 1231-1237. DOI:10.1177/0278364913491297 .
[18]
MAHJOURIANR, WICKEM, ANGELOVAA. Unsupervised learning of depth and ego-motion from monocular video using 3D geometric constraints[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 5667-5675. DOI:10.1109/CVPR.2018.00594 .
[19]
ZOUY L, LUOZ L, HUANGJ B. DF-net: Unsupervised joint learning of depth and flow using cross-task consistency[C]//Computer Vision—ECCV 2018(LNCS 11209). Berlin:Springer, 2018: 38-55. DOI:10.1007/978-3-030-01228-1_3 .
[20]
WANGC Y, BUENAPOSADAJ M, ZHUR, et al. Learning depth from monocular videos using direct methods[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 2022-2030. DOI:10.1109/CVPR.2018.00216 .
[21]
RANJANA, JAMPANIV, BALLESL, et al. Competitive collaboration: Joint unsupervised learning of depth, camera motion, optical flow and motion segmentation[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2019: 12232-12241. DOI:10.1109/CVPR.2019.01252 .
[22]
SPENCERJ, BOWDENR, HADFIELDS. DeFeat-net: General monocular depth via simultaneous unsupervised representation learning[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2020: 14390-14401. DOI:10.1109/CVPR42600.2020.01441 .
[23]
LIH H, GORDONA, ZHAOH, et al. Unsupervised monocular depth learning in dynamic scenes[EB/OL]. 2020: arXiv: 2010.16404.
[24]
ZHANH Y, GARGR, WEERASEKERAC S, et al. Unsupervised learning of monocular depth estimation and visual odometry with deep feature reconstruction[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 340-349. DOI:10.1109/CVPR.2018.00043 .
[25]
MUR-ARTALR, TARDÓSJ D. ORB-SLAM2: An open-source SLAM system for monocular, stereo, and RGB-D cameras[J]. IEEE Transactions on Robotics, 2017, 33(5): 1255-1262. DOI:10.1109/TRO.2017.2705103 .
[26]
SHENT W, LUOZ X, ZHOUL, et al. Beyond photometric loss for self-supervised ego-motion estimation[C]//2019 International Conference on Robotics and Automation (ICRA). New York: IEEE Press, 2019: 6359-6365. DOI:10.1109/ICRA.2019.8793479 .
[27]
BIANJ W, ZHANH Y, WANGN Y, et al. Unsupervised scale-consistent depth learning from video[J]. International Journal of Computer Vision, 2021, 129(9): 2548-2564. DOI:10.1007/s11263-021-01484-6 .