1.School of Computer Science,Wuhan University,Wuhan 430072,Hubei,China
2.Wuhan Maritime Communication Research Institute,Wuhan 430205,Hubei,China
Show less
文章历史+
Received
Published
2021-05-25
2023-02-24
Issue Date
2026-07-23
PDF (2641K)
摘要
在RGB-D显著性检测视觉任务中,RGB彩色模态和深度模态的信息均被视为十分重要的特征线索。但现有的RGB-D显著性检测模型无法高效执行多尺度特征的交互和多模态特征的融合,因此在真实的开放场景下表现欠佳。针对上述问题,提出了一种基于协同注意力(synergistic attention)机制的RGB-D显著性检测算法模型(SANet),并引入多模态学习中通用的引导与教导策略(guidance and teaching strategy)。在编码器进行多尺度特征提取的阶段中进行隐式引导(implicit guidance),在解码器进行特征融合时进行显式的教导(explicit teaching),实现了编码、解码的分阶段学习。在4个显著性检测评测数据集上进行的综合实验表明,该算法在4个评测指标上均优于已有的18个前沿RGB-D显著性检测模型。
Abstract
RGB information and depth information are both crucial feature cues in RGB-D saliency detection. However, the current RGB-D saliency detection models are unable to handle the interaction of multi-scale features and the fusion of multi-modal features efficiently, so they are limited in understanding natural scenes. To alleviate such shortcomings, we propose a synergistic attention network for RGB-D saliency detection (SANet), and introduce a general multi-modal learning strategy, termed Guidance and Teaching Strategy. The implicit guidance strategy works in the multi-scale feature extraction stage of the encoder, while the explicit teaching strategy works in the feature fusion stage of the decoder, so as to realize the phased learning process of encoding and decoding stages. Extensive experiments are conducted on four saliency detection benchmarks, and the results verify the proposed method performs better than the 18 cutting-edge RGB-D saliency detection models in terms of four metrics.
GUOY C, YUANH J, WUP. Image saliency detection based on local and regional features[J]. Acta Automatica Sinica, 2013, 39(8):1214-1224. DOI: 10.3724/SP.J.1004.2013.01214 (Ch ).
[3]
ZHUC B, LIG, WANGW M, et al. An innovative salient object detection using center-dark channel prior [C]// Proceedings of the IEEE International Conference on Computer Vision Workshops. Piscataway: IEEE Press, 2017: 1509-1515. DOI: 10.1109/ICCVW.2017.178 .
[4]
LIANGF F, DUANL J, MAW, et al. Stereoscopic saliency model using contrast and depth-guided-background prior [J]. Neurocomputing, 2018, 275: 2227-2238. DOI: 10.1016/j.neucom.2017.10.052 .
[5]
PIAOY R, RONGZ K, ZHANGM, et al. A2dele: Adaptive and attentive depth distiller for efficient RGB-D salient object detection [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2020: 9057-9066. DOI:10.1109/CVPR42600.2020.00908 .
[6]
ZHANGM, RENW S, PIAOY R, et al. Select, supplement and focus for RGB-D saliency detection [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2020: 3469-3478. DOI:10.1109/CVPR42600.2020.00353 .
[7]
ZHAOX Q, ZHANGL H, PANGY W, et al. A single stream network for robust and real-time RGB-D salient object detection [C]// European Conference on Computer Vision. Cham: Springer International Publishing, 2020: 646-662. DOI: 10.1007/978-3-030-58542-6_39 .
JIAOY X, WANGX, CHOUY C, et al. Guidance and teaching network for video salient object detection[C]//2021 IEEE International Conference on Image Processing. New York: IEEE Press, 2021: 2199-2203. DOI:10.1109/ICIP42928.2021.9506492 .
HEK M, ZHANGX Y, RENS Q, et al. Deep residual learning for image recognition[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2016: 770-778. DOI:10.1109/CVPR.2016.90 .
[13]
LUJ S, YANGJ W, BATRAD, et al. Hierarchical question-image co-attention for visual question answering[J]. Advances in Neural Information Processing Systems, 2016, 29: 289-297.
DHINGRAB, LIUH X, YANGZ L, et al. Gated-attention readers for text comprehension [C]// Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics. Stroudsburg: Association for Computational Linguistics, 2017: 1832-1846. DOI: 10.18653/v1/p17-1168 .
[16]
LAIT, TRANQ H, BUI T, et al. A Gated Self-Attention Memory Network for Answer Selection[EB/OL]. [2019-09-13]. DOI: 10.18653/v1/D19-1610 .
[17]
LIUS T, HUANGD, WANGY H. Receptive field block net for accurate and fast object detection [C]// Proceedings of the European Conference on Computer Vision (ECCV). Berlin: Springer, 2018: 385-400. DOI: 10.1007/978-3-030-01252-6_24 .
LIN Y, YEJ W, JIY, et al. Saliency detection on light field[C]//2014 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2014: 2806-2813. DOI:10.1109/CVPR.2014.359 .
[20]
PENGH W, LIB, XIONGW H, et al. RGBD salient object detection: A benchmark and algorithms[C]//Computer Vision―ECCV 2014. Berlin: Springer, 2014: 92-109. DOI:10.1007/978-3-319-10578-9_7 .
[21]
FAND P, LINZ, ZHANGZ, et al. Rethinking RGB-D salient object detection: Models, data sets, and large-scale benchmarks [J]. IEEE Transactions on Neural Networks and Learning Systems, 2021, 32(5): 2075-2089. DOI:10.1109/TNNLS.2020.2996406 .
[22]
ZHAOJ X, CAOY, FAND P, et al. Contrast prior and fluid pyramid integration for RGBD salient object detection[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2019: 3922-3931. DOI:10.1109/CVPR.2019.00405 .
[23]
CHENH, LIY F. Three-stream attention-aware network for RGB-D salient object detection [J]. IEEE Transactions on Image Processing, 2019, 28(6): 2825-2835. DOI: 10.1109/TIP.2019.2891104 .
[24]
ZHUC B, CAIX, HUANGK, et al. PDNet: Prior-model guided depth-enhanced network for salient object detection[C]//2019 IEEE International Conference on Multimedia and Expo. New York: IEEE Press, 2019: 199-204. DOI:10.1109/ICME.2019.00042 .
[25]
JUR, GEL, GENGW J, et al. Depth saliency based on anisotropic center-surround difference[C]//2014 IEEE International Conference on Image Processing. New York: IEEE Press, 2014: 1115-1119. DOI:10.1109/ICIP.2014.7025222 .
[26]
DENGJ, DONGW, SOCHERR, et al. ImageNet: A large-scale hierarchical image database[C]//2009 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2009: 248-255. DOI:10.1109/CVPR.2009.5206848 .
[27]
HEK M, ZHANGX Y, RENS Q, et al. Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification[C]//2015 IEEE International Conference on Computer Vision. New York: IEEE Press, 2015: 1026-1034. DOI:10.1109/ICCV.2015.123 .
[28]
PIAOY R, JIW, LIJ J, et al. Depth-induced multi-scale recurrent attention network for saliency detection[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV). New York: IEEE Press, 2019: 7253-7262. DOI:10.1109/ICCV.2019.00735 .
[29]
FAND P, CHENGM M, LIUY, et al. Structure-measure: A new way to evaluate foreground maps[C]//2017 IEEE International Conference on Computer Vision. New York: IEEE Press, 2017: 4558-4567. DOI:10.1109/ICCV.2017.487 .
[30]
ACHANTAR, HEMAMIS, ESTRADAF, et al. Frequency-tuned salient region detection[C]//2009 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2009: 1597-1604. DOI:10.1109/CVPR.2009.5206596 .
[31]
FAND P, GONGC, CAOY, et al. Enhanced-alignment measure for binary foreground map evaluation[C]//Proceedings of the 27th International Joint Conference on Artificial Intelligence. California: International Joint Conferences on Artificial Intelligence Organization, 2018: 698-704. DOI:10.24963/ijcai.2018/97 .
[32]
PERAZZIF, KRÄHENBÜHLP, PRITCHY, et al. Saliency filters: Contrast based filtering for salient region detection[C]//2012 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2012: 733-740. DOI:10.1109/CVPR.2012.6247743 .
[33]
FENGD, BARNESN, YOUS D, et al. Local background enclosure for RGB-D salient object detection [C]//2016 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2016: 2343-2350. DOI:10.1109/CVPR.2016.257 .
[34]
GUOJ F, RENT W, BEIJ. Salient object detection for RGB-D image via saliency evolution[C]//2016 IEEE International Conference on Multimedia and Expo. New York: IEEE Press, 2016: 1-6. DOI:10.1109/ICME.2016.7552907 .
[35]
CONGR M, LEIJ J, ZHANGC Q, et al. Saliency detection for stereoscopic images based on depth confidence analysis and multiple cues fusion[J]. IEEE Signal Processing Letters, 2016, 23(6): 819-823. DOI:10.1109/LSP.2016.2557347 .
[36]
QUL Q, HES F, ZHANGJ W, et al. RGBD salient object detection via deep fusion[J]. IEEE Transactions on Image Processing, 2017, 26(5): 2274-2285. DOI:10.1109/TIP.2017.2682981 .
[37]
HANJ W, CHENH, LIUN, et al. CNNs-based RGB-D saliency detection via cross-view transfer and multiview fusion[J]. IEEE Transactions on Cybernetics, 2018, 48(11): 3171-3183. DOI:10.1109/TCYB.2017.2761775 .
[38]
SONGH K, LIUZ, DUH, et al. Depth-aware salient object detection and segmentation via multiscale discriminative saliency fusion and bootstrap learning [J]. IEEE Transactions on Image Processing, 2017, 26(9): 4204-4216. DOI: 10.1109/TIP.2017.2711277 .
[39]
CHENH, LIY F. Progressively complementarity-aware fusion network for RGB-D salient object detection[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2018: 3051-3060. DOI:10.1109/CVPR.2018.00322 .
[40]
WANGN N, GONGX J. Adaptive fusion for RGB-D salient object detection [J]. IEEE Access, 2019, 7: 55277-55284. DOI: 10.1109/ACCESS.2019.2913107 .
[41]
CHENH, LIY F, SUD. Multi-modal fusion network with multi-scale multi-path and cross-modal interactions for RGB-D salient object detection [J]. Pattern Recognition, 2019, 86: 376-385. DOI: 10.1016/j.patcog.2018.08.007 .