In order to solve the problems of fixed feature dimensions, high computational complexity, and serious category imbalance in multimodal feature fusion in the current Salient Object Detection (SOD), and the deficiencies of the existing methods in terms of insufficient context modelling, loss of detailed features, and cross-modal coupling stiffness lead to the limited detection performance, this paper proposes a novel bimodal gated cross-modal attention salient target detection method, ECCLNet, to improve detection accuracy and efficiency. The Cross-modal Enhancement Module (CEM) is adopted to achieve adaptive interaction between RGB and depth/thermal infrared modal features, and the Cross-Level Refinement Module (CLRM) is designed to combine deformable convolution and dynamic weighting strategies to achieve accurate alignment and detection of multi-scale features. Meanwhile, for the problem of category imbalance, a Weighted Binary Cross-Entropy (WBCE) loss function based on the square root balancing strategy is proposed to stabilise the training and improve the boundary prediction accuracy. Comparative experiments are carried out on standard RGB-D and RGB-T salient object detection data-sets of NLPR, SIP and VT series. The results indicate that the proposed ECCLNet outperforms all competing algorithms in both detection accuracy and inference efficiency, and can offer technical support for studies related to multi-modal visual perception and intelligent object detection under complex scenarios.
KimJ H, LeeJ. Layered non-photorealistic rendering with anisotropic depth-of-field filtering[J]. Multimedia Tools and Applications, 2020, 79(1): 1291-1309.
[2]
LeT N, SugimotoA. Video salient object detection using spatiotemporal deep features[J]. IEEE Transactions on Image Processing, 2018, 27(10): 5002-5015.
[3]
WeiS K, LiaoL X, LiJ, et al. Saliency inside: learning attentive CNNs for content-based image retrieval[J]. IEEE Transactions on Image Processing, 2019, 28(9): 4580-4593.
[4]
YueH H, GuoJ C, YinX J, et al. Salient object detection in low-light images via functional optimization-inspired feature polishing[J]. Knowledge-Based Systems, 2022, 257: 109938.
[5]
WangX H, LiuZ B, LiesaputraV, et al. Feature specific progressive improvement for salient object detection[J]. Pattern Recognition, 2024, 147: 110085.
[6]
ZhouW J, GuoQ L, LeiJ S, et al. ECFFNet: effective and consistent feature fusion network for RGB-T salient object detection[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2022, 32(3): 1224-1235.
[7]
FanD P, ZhaiY J, BorjiA, et al. BBS-Net: RGB-D salient object detection with a bifurcated backbone strategy network[C]//Computer Vision-ECCV 2020. Cham: Springer, 2020:275-292.
[8]
ZhaoX Q, PangY W, ZhangL H, et al. Suppress and balance: a simple gated network for salient object detection[C]//Computer Vision-ECCV 2020. Cham: Springer, 2020: 35-51.
[9]
ZhaoJ X, LiuJ J, FanD P, et al. EGNet: edge guidance network for salient object detection[C]//2019 IEEE/CVF International Conference on Computer Vision. October 27-November 2, 2019. Seoul, Korea. IEEE, 2019: 8778-8787.
[10]
BaoL X, ZhouX F, LuX K, et al. Quality-aware selective fusion network for V-D-T salient object detection[J]. IEEE Transactions on Image Processing, 2024, 33: 3212-3226.
[11]
LiangB C, LuoH L. MEANet: an effective and lightweight solution for salient object detection in optical remote sensing images[J]. Expert Systems with Applications, 2024, 238: 1217785.
[12]
HaoC, YuZ T, LiuX, et al. A simple yet effective network based on vision transformer for camouflaged object and salient object detection[J]. IEEE Transactions on Image Processing, 2025, 34: 608-622.
[13]
YuanG J, SongJ T, LiJ J. IF-USOD: multimodal information fusion interactive feature enhancement architecture for underwater salient object detection[J]. Information Fusion, 2025, 117: 102806.
[14]
YangQ N, ZhengJ H, ChenJ. Multilevel diverse feature aggregation network for salient object detection[J]. Neurocomputing, 2025, 628: 129648.
[15]
XuM Y, ZhouZ P, XuH B, et al. CP-net: contour-perturbed reconstruction network for self-supervised point cloud learning[J]. IEEE Transactions on Multimedia, 2024, 26: 8799-8810.
[16]
FuK R, FanD P, JiG P, et al. JL-DCF: Joint learning and densely-cooperative fusion framework for RGB-D salient object detection[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2020: 3049-3059.
[17]
PengD G, ZhouW Y, PanJ Z, et al. MSEDNet: multi-scale fusion and edge-supervised network for RGB-T salient object detection[J]. Neural Networks, 2024, 171: 410-422.
[18]
ChenJ, ZhangH Y, GongM M, et al. Collaborative compensative transformer network for salient object detection[J]. Pattern Recognition, 2024, 154: 110600.
[19]
LiuY, LiC X, DongX H, et al. Seamless detection: unifying salient object detection and camouflaged object detection[J]. Expert Systems with Applications, 2025, 274: 126912.
ChenHong, LiHongxu, Jinhaibo. Intrusion detection method based on multi-scale convolution and dual attention mechanism[J].Journal of Liaoning Technical University (Natural Science),2024,43(1):93-100.
ChenZhanguo, ChenZhenjun, XueChenxia, et al. A complex scene depth estimation network based on monocular camera[J].Journal of Liaoning Technical University (Natural Science),2025,44(4):505-512.