Self-supervised monocular depth estimation methods trained on sequences of monocular images have received considerable attention in recent years by using the photometric consistency loss between adjacent frames instead of depth labels as the supervisory signal for network training. The photometric consistency constraint follows the static world assumption, but the moving objects in the monocular image sequence violate this assumption, which affects the camera pose estimation accuracy and the calculation accuracy of the photometric loss function during the self-supervised training process. By detecting and removing the moving target area, the camera pose decoupled from the target motion can be obtained, and the influence of the moving target area on the calculation accuracy of the photometric loss can be discarded. To this end, this paper proposes a self-supervised monocular depth estimation network based on semantic assistance and depth temporal consistency constraints. First, an offline instance segmentation network is used to detect dynamic category objects that may violate the static world assumption, and the corresponding region input pose network is removed to obtain a camera pose decoupled from object motion. Secondly, based on semantic consistency and photometric consistency constraints, the motion status of dynamic category targets is detected so that the photometric loss in the moving area does not affect the iterative update of network parameters.Finally, depth temporal consistency constraints are imposed in non-motion areas, and the estimated depth value of the current frame is explicitly aligned with the projected depth value of adjacent frames to further refine the depth prediction results. Experiments on the KITTI, DDAD and KITTI Odometry datasets verify that the proposed method has better performance than previous self-supervised monocular depth estimation methods.
对于KITTI数据集,输入图像的分辨率调整为832像素×256像素用于训练.此外,在训练阶段会应用数据增强来提高网络的鲁棒性.数据增强策略包括随机水平翻转和缩放,增广概率均为0.5.训练样本长度设置为3,即当前帧、当前帧的前一帧和后一帧.使用Adam优化器训练网络,批处理大小和学习率分别设置为8和1e-4,学习总轮次为50.损失函数中的超参数设置为w1=1,w2=0.5,w3=0.1.深度评测范围为0~80 m.
LAGAH, JOSPINL V, BOUSSAIDF,et al .A survey on deep learning techniques for stereo-based depth estimation[J].IEEE Trans Pattern Anal Mach Intell,2022,44(4):1738-1764.
[2]
FURUKAWAY, HERNÁNDEZC .Multi-view stereo:a tutorial[J].Foundations and Trends in Computer Graphics and Vision,2015,9(1/2):1-148.
[3]
MINGY, MENGX Y, FANC X, et al .Deep learning for monocular depth estimation:a review[J].Neurocomputing, 2021,438: 14-33.
[4]
ZHANGR, TSAIP S, CRYERJ E, et al .Shape-from-shading:a survey[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,1999,21(8):690-706.
[5]
TAOM W, HADAPS, MALIKJ, et al .Depth from combining defocus and correspondence using light-field cameras[C]//2013 IEEE International Conference on Computer Vision.Sydney,NSW,Australia. IEEE,2013:673-680.
[6]
SAXENAA, CHUNGS H, NGA Y .Learning depth from single monocular images[C]//Proceedings of the 18th International Conference on Neural Information Processing Systems. Vancouver, British Columbia,Canada. ACM,2005:1161-1168.
[7]
KARSCHK, LIUC, KANGS B .Depth transfer:depth extraction from video using non-parametric sampling[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2014, 36(11):2144-2158.
[8]
EIGEND, PUHRSCHC, FERGUSR .Depth map prediction from a single image using a multi-scale deep network[C]//Proceedings of the 27th International Conference on Neural Information Processing Systems-Volume 2. Montreal,Canada. ACM,2014:2366-2374.
[9]
FUH, GONGM M, WANGC H, et al .Deep ordinal regression network for monocular depth estimation[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Salt Lake City,UT,USA. IEEE,2018:2002-2011.
[10]
GARGR, VIJAY KUMARB G, CARNEIROG, et al .Unsupervised CNN for single view depth estimation:geometry to the rescue[M]//LEIBE B,MATAS J,SEBE N,et al,eds.Computer Vision – ECCV 2016.Cham:Springer International Publishing,2016:740-756.
[11]
ZHOUT H, BROWNM, SNAVELYN, et al .Unsupervised learning of depth and ego-motion from video[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).Honolulu,HI,USA. IEEE,2017:6612-6619.
[12]
LIUL, ZHAIG Y, YEW L, et al .Unsupervised learning of scene flow estimation fusing with local rigidity[C]//IJCAI, 2019:876-882.
[13]
KLINGNERM, TERMÖHLENJ A, MIKOLAJCZYKJ, et al .Self-supervised monocular depth estimation:solving the dynamic object problem by semantic guidance[C]//European Conference on Computer Vision.Cham:Springer,2020:582-600.
[14]
ZHANGH K, LIY, CAOY, et al .Exploiting temporal consistency for real-time video depth estimation[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV).Seoul,Korea (South). IEEE,2019:1725-1734.
[15]
LIS Y, LUOY, ZHUY, et al .Enforcing temporal consistency in video depth estimation[C]//2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW).Montreal,BC,Canada. IEEE,2021:1145-1154.
[16]
KOPFJ, RONGX J, HUANGJ B .Robust consistent video depth estimation[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Nashville,TN,USA. IEEE,2021:1611-1621.
[17]
JADERBERGM, SIMONYANK, ZISSERMANA, et al. Spatial transformer networks[EB/OL].
[18]
HEK M, GKIOXARIG, DOLLÁRP, et al .Mask R-CNN[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2020,42(2):386-397.
[19]
WANGZ, BOVIKA C, SHEIKHH R, et al .Image quality assessment:from error visibility to structural similarity[J].IEEE Transactions on Image Processing,2004,13(4):600-612.
[20]
GODARDC, AODHA OMAC, FIRMANM, et al .Digging into self-supervised monocular depth estimation[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). Seoul,Korea (South). IEEE,2019:3827-3837.
[21]
SAUNDERSK, VOGIATZISG, MANSOL J .Dyna-DM:dynamic object-aware self-supervised monocular depth maps[C]//2023 IEEE International Conference on Autonomous Robot Systems and Competitions (ICARSC).Tomar,Portugal. IEEE,2023:10-16.
[22]
GEIGERA, LENZP, STILLERC, et al .Vision meets robotics:the KITTI dataset[J]. International Journal of Robotics Research,2013,32(11): 1231-1237.
[23]
GEIGERA, LENZP, URTASUNR .Are we ready for autonomous driving?The KITTI vision benchmark suite[C]//2012 IEEE Conference on Computer Vision and Pattern Recognition.Providence,RI,USA. IEEE, 2012: 3354-3361.
[24]
CHENP Y, LIUA H, LIUY C, et al .Towards scene understanding:unsupervised monocular depth estimation with semantic-aware representation[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Long Beach,CA,USA. IEEE,2019:2619-2627.
[25]
PILLAIS, AMBRUŞR, GAIDONA. SuperDepth:self-supervised,super-resolved monocular depth estimation[C]//2019 International Conference on Robotics and Automation (ICRA).Montreal,QC,Canada. IEEE, 2019: 9250-9256.
ZHOUD K, TIANJ, YANGX .Unsupervised monocular image depth estimation based on the prediction of local plane parameters[J]. Journal of Image and Graphics,2021,26(1):165-175.(in Chinese)
[28]
MAHJOURIANR, WICKEM, ANGELOVAA .Unsupervised learning of depth and ego-motion from monocular video using 3D geometric constraints[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Salt Lake City,UT,USA. IEEE,2018:5667-5675.
[29]
WANGC Y, BUENAPOSADAJ M, ZHUR, et al .Learning depth from monocular videos using direct methods[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Salt Lake City,UT,USA. IEEE,2018:2022-2030.
[30]
CHENS, PUZ D, FANX, et al .Fixing defect of photometric loss for self-supervised monocular depth estimation[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2022,32(3):1328-1338.
[31]
BIANJ W, ZHANH Y, WANGN Y, et al .Unsupervised scale-consistent depth learning from video[J].International Journal of Computer Vision,2021,129(9):2548-2564.
[32]
ZHANGY R, GONGM G, LIJ Z, et al .Self-supervised monocular depth estimation with multiscale perception[J].IEEE Transactions on Image Processing:a Publication of the IEEE Signal Processing Society,2022,31:3251-3266.
[33]
ZHANH Y, GARGR, WEERASEKERAC S, et al .Unsupervised learning of monocular depth estimation and visual odometry with deep feature reconstruction[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Salt Lake City,UT,USA. IEEE,2018:340-349.
[34]
LIR H, WANGS, LONGZ Q, et al .UnDeepVO:monocular visual odometry through unsupervised deep learning[C]//2018 IEEE International Conference on Robotics and Automation (ICRA).Brisbane,QLD,Australia. IEEE,2018:7286-7291.
[35]
YINZ C, SHIJ P .GeoNet:unsupervised learning of dense depth,optical flow and camera pose[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Salt Lake City,UT,USA. IEEE,2018:1983-1992.
[36]
SHENT W, LUOZ X, ZHOUL, et al .Beyond photometric loss for self-supervised ego-motion estimation[C]//2019 International Conference on Robotics and Automation (ICRA).Montreal,QC,Canada. IEEE,2019:6359-6365.
基金资助
国家自然科学基金资助项目(62171184)
国家自然科学基金资助项目(62273139)
国家自然科学基金资助项目(62106072)
National Natural ScienceFoundation of China(62171184)
National Natural ScienceFoundation of China(62273139)
National Natural ScienceFoundation of China(62106072)
国家自然科学基金区域联合重点项目(U23A20385)
Joint Funds of the National Natural Science Foundation of China(U23A20385)
国防预研项目(JCY2021206B015)
National Defense Pre-research Foundation(JCY2021206B015)