To solve the problems of the existing target tracking algorithms, such as inability to extract deep-level features, failure to fully exploit cross-modal information, and weak representation of target features, a feature fusion shift Siamese network for RGB-T target tracking is proposed. First, a target tracking framework based on the visible modal SiameseRPN++ is designed to extend the infrared modal branch, in order to obtain a multimodal target tracking framework. Moreover, the improved ResNet50 network with adjusted stride as a feature extraction network enables the acquisition of deep-level features of the target. Subsequently, a multimodal feature interactive learning module (FIM) is designed to leverage the discriminative information from one modality to guide the learning process of target appearance features in the other modality. By mining the cross-modal information within the feature space and channels, the module enhances the network’s attention towards foreground information. Thereafter, a multimode feature fusion module (FAM) is designed, which calculates the degree of feature fusion between the input visible light image and the infrared image, enabling spatial fusion of significant features from different modalities to effectively eliminate redundant information and reconstructing multimodal images by employing a cascade fusion strategy. Finally, a feature space shift module (FSM) is designed, which divides the feature maps of the infrared modal branches and shifts them in four different directions to enhance the edge representation of the heat source target. Extensive experiments on two RGB-T datasets thoroughly validate the effectiveness of the proposed algorithm, while ablation experiments demonstrate the superiority of each designed module.
DANELLJANM, BHATG, KHANF S,et al .ATOM:accurate tracking by overlap maximization[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Long Beach,CA,USA. IEEE,2019:4660–4669.
[2]
BHATG, DANELLJANM, VAN GOOLL,et al. Learning discriminative model prediction for tracking[C]//2019 IEEE/CVF International Conference on Computer Vision (ICCV).Seoul,Korea (South). IEEE, 2019: 6182-6191.
[3]
NAM H, HANB .Learning multi-domain convolutional neural networks for visual tracking[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas,NV,USA. IEEE,2016:4293-4302.
[4]
LANX Y, YEM, SHAOR,et al .Learning modality-consistency feature templates:a robust RGB-infrared tracking system[J].IEEE Transactions on Industrial Electronics,2019,66(12):9887-9897.
[5]
LIC L, ZHUC L, HUANGY,et al .Cross-modal ranking with soft consistency and noisy labels for robust RGB-T tracking[C]// Computer Vision – ECCV 2018.Cham:Springer International Publishing,2018:831-847.
[6]
LIC L, CHENGH, HUS Y,et al .Learning collaborative sparse representation for grayscale-thermal tracking[J]. IEEE Transactions on Image Processing,2016,25(12):5743-5756.
[7]
LIC L, SUNX, WANGX,et al. Grayscale-thermal object tracking via multitask Laplacian sparse representation[J].IEEE Transactions on Systems,Man,and Cybernetics:Systems, 2017,47(4): 673-681.
[8]
GUOC, YANGD D, LIC,et al .Dual Siamese network for RGBT tracking via fusing predicted position maps[J]. The Visual Computer, 2022, 38(7): 2555-2567.
[9]
GUOC Y, XIAOL .High speed and robust RGB-thermal tracking via dual attentive stream Siamese network[C]//IGARSS 2022—2022 IEEE International Geoscience and Remote Sensing Symposium. Kuala Lumpur,Malaysia. IEEE,2022:803-806.
[10]
LIC L, LUA D, ZHENGA H,et al .Multi-adapter RGBT tracking[C]//2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW). Seoul,Korea (South). IEEE,2019.
HEK M, ZHANGX Y, RENS Q,et al .Deep residual learning for image recognition[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas,NV,USA. IEEE,2016:770-778.
[13]
ZHANGT L, LIUX R, ZHANGQ,et al .SiamCDA:complementarity- and distractor-aware RGB-T tracking based on Siamese network[J].IEEE Transactions on Circuits and Systems for Video Technology, 2022, 32(3): 1403-1417.
[14]
LIY D, LAIH C, WANGL J,et al .Multibranch adaptive fusion network for RGBT tracking[J].IEEE Sensors Journal, 2022, 22(7):7084-7093.
[15]
TANGZ Y, XUT Y, LIH,et al .Exploring fusion strategies for accurate RGBT visual object tracking[EB/OL].2022:2201.08673.
[16]
LIB, WUW, WANGQ,et al .SiamRPN:evolution of Siamese visual tracking with very deep networks[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Long Beach,CA,USA. IEEE,2019:4277-4286.
[17]
GIRSHICKR .Fast R-CNN[C]//2015 IEEE International Conference on Computer Vision (ICCV). Santiago,Chile. IEEE, 2015: 1440-1448.
[18]
RUSSAKOVSKYO, DENGJ, SUH,et al .ImageNet large scale visual recognition challenge[J]. International Journal of Computer Vision, 2015, 115(3): 211-252.
[19]
REALE, SHLENSJ, MAZZOCCHIS,et al.YouTube-BoundingBoxes:a large high-precision human-annotated data set for object detection in video[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu,HI,USA. IEEE,2017:7464-7473.
[20]
LINT Y, MAIREM, BELONGIES, et al. Microsoft coco: common objects in context [C]// Computer Vision-ECCV 2014. Zurich, Switzerland. Springer, 2014: 740-755.
[21]
LIC L, LIANGX Y, LUY J,et al .RGB-T object tracking:benchmark and baseline[J].Pattern Recognition,2019,96:106977.
[22]
LIC L, XUEW L, JIAY Q,et al .LasHeR:a large-scale high-diversity benchmark for RGBT tracking[J].IEEE Transactions on Image Processing,2021,31:392-404.
[23]
GAOY, LIC L, ZHUY B,et al. Deep adaptive fusion network for high performance RGBT tracking[C]//2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW). Seoul,Korea (South). IEEE, 2019: 91-99.
[24]
ZHANGZ P, PENGH W .Deeper and wider Siamese networks for real-time visual tracking[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Long Beach,CA,USA. IEEE, 2019: 4591-4600.
[25]
ZHANGH, ZHANGL, ZHUOL,et al. Object tracking in RGB-T videos using modal-aware attention network and competitive learning[J].Sensors,2020,20(2):393.
[26]
LIC L, ZHAON, LUY J,et al .Weighted sparse representation regularized graph learning for RGB-T object tracking[C]//Proceedings of the 25th ACM International Conference on Multimedia. Mountain View, California, USA. ACM,2017:1856-1864.
[27]
KIMH U, LEED Y, SIMJ Y,et al .SOWP:spatially ordered and weighted patch descriptor for visual tracking[C]//2015 IEEE International Conference on Computer Vision (ICCV). Santiago,Chile. IEEE, 2015: 3011-3019.
[28]
DANELLJANM, BHATG, KHANF S,et al .ECO:efficient convolution operators for tracking[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu,HI,USA. IEEE,2017:6931-6939.