The application of multi-object tracking (MOT) technology opens up new possibilities for team sports video monitoring and analysis, enabling real-time tracking of multiple athletes and supporting multidimensional analysis and understanding of game dynamics. However, in complex team sports scenarios, issues such as mutual occlusion between athletes, rapid movements, and frequent changes in target identities may potentially degrade tracking performance. To address these challenges, an end-to-end deep learning MOT framework based on VisionTransformer was proposed, which mainly consisted of two parts: detection network and memory network. The detection network comprised a convolutional neural network (CNN) backbone, Vision Transformer encoder and decoder. The ResNet50 was adopted as a feature extractor, and the traditional feed-forward neural network (FFN) layer was replaced by a local attention (LA) module to obtain more comprehensive feature representations through the combination of global attention and local convolution. The memory network consisted of a memory encoding module and spatio-temporal memory decoder. The memory encoding module was responsible for aggregating the target embedding information, in which the short-term cross attention (CA) module focused on the immediate states, while the long-term CA module explored the significant features covered by memory over time spans, captured dependencies and associations over long time intervals to effectively preserve temporal context information of tracked objects. The spatio-temporal memory decoder integrated encoded frame embeddings, candidate embeddings and trajectory embeddings to address multi-object detection and identity association in MOT. The spatio-temporal memory mechanism efficiently retained observed historical states of targets and combined with an attention mechanism, accurately predicted target states. Experimental results demonstrate that the proposed framework achieves 75.7% HOTA and 98.5% MOTA on the team sports video public dataset SportsMOT, outperforming other state-of-the-art MOT methods. Additionally, the proposed framework achieves optimal or near-optimal performance on multiple metrics on the generalized public datasets MOT17 and MOT20, further validating the effectiveness and robustness of the proposed framework.
PALS K, PRAMANIKA, MAITIJ, et al. Deep learning in multi-object detection and tracking: state of the art[J]. Applied Intelligence, 2021, 51(9): 6400-6429.
ZHOUXue, LIANGChao, HEJunyang, et al. A survey on one-shot multi-object tracking algorithm[J]. Journal of University of Electronic Science and Technology of China, 2022, 51(5): 728-736. (in Chinese)
[4]
YAOR, LING, XIAS, et al. Video object segmentation and tracking: A survey[J]. ACM Transactions on Intelligent Systems and Technology (TIST), 2020, 11(4): 1-47.
YANKang, ZENGFengcai, HENing, et al. JDE Multi-Object Tracking Method with Attention Mechanism[J]. Computer Engineering and Applications, 2022, 58(21): 189-196. (in Chinese)
[7]
BEWLEYA, GEZ, OTT L, et al. Simple online and realtime tracking[C]//2016 IEEE International Conference on Image Processing (ICIP). IEEE, 2016: 3464-3468.
[8]
DUY, ZHAOZ, SONGY, et al. Strongsort: Make deepsort great again[J]. IEEE Transactions on Multimedia, 2023, 25: 8725-8737.
GuiE, WANGYongxiong. Multi-candidate association online multi-target tracking based on R-FCN framework[J]. Opto-Electronic Engineering, 2020, 47(1): 29-37. (in Chinese)
[11]
XUJ, CAOY, ZHANGZ, et al. Spatial-temporal relation networks for multi-object tracking[C]//2019 Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, 2019: 3987-3997.
[12]
ZHOUX, KOLTUNV, KRÄHENBÜHLP. Tracking objects as points[C]//2020 European Conference on Computer Vision(ECCV). Springer Science, 2020 : 474-490.
[13]
ZHANGY, WANGC, WANGX, et al. FairMOT: On the fairness of detection and re-identification in multiple object tracking[J]. International Journal of Computer Vision, 2021, 129(11): 3069-3087.
LIQingge, YANGXiaogang, LURuitao, et al. Transformer in computer vision: A survey[J]. Journal of Chinese Computer Systems, 2023, 44(4): 850-861. (in Chinese)
[16]
MEINHARDTT, KIRILLOVA, LEAL-TAIXÉL, et al. Trackformer: Multi-object tracking with transformers[C]//2022 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR). IEEE, 2022: 8834-8844.
[17]
XUY, BANY, DELORMEG, et al. TransCenter: Transformers with dense representations for multiple-object tracking[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(6): 7820-7835.
[18]
CHUP, WANGJ, YOUQ, et al. TransMOT: Spatial-temporal graph transformer for multiple object tracking[C]//2023 Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). IEEE, 2023: 4859-4869.
WANGNing, XiMao, ZHOUWengang, et al. Recent advance in deep visual object tracking[J]. Journal of University of Science and Technology of China, 2021, 51(4): 335-344. (in Chinese)
[21]
ARKINE, YADIKARN, XUX, et al. A survey: Object detection methods from CNN to transformer[J]. Multimedia Tools and Applications, 2023, 82(14): 21353-21383.
ZHANGTao, ZHANGXiaoli, RENYan. Monocular image depth estimation based on the fusion of transformer and CNN[J]. Journal of Harbin University of Science and Technology, 2022, 27(6): 88-94. (in Chinese)
[24]
LIK, YUR, WANGZ, et al. Locality guidance for improving vision transformers on tiny datasets[C]//2022 European Conference on Computer Vision (ECCV), 2022: 110-127.
[25]
SONGC H, YOONJ, CHOIS, et al. Boosting vision transformers for image retrieval[C]//2023 Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). IEEE, 2023: 107-117.
[26]
CAIJ, XUM, LIW, et al. MeMOT: Multi-object tracking with memory[C]//2022 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2022: 8080-8090.
[27]
LINT Y, GOYALP, GIRSHICKR, et al. Focal loss for dense object detection[C]//2017 Proceedings of the IEEE International Conference on Computer vision (ICCV). IEEE, 2017: 2899-3007.
[28]
REZATOFIGHIH, TSOIN, GWAKJ Y, et al. Generalized intersection over union: A metric and a loss for bounding box regression[C]//2019 Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2019: 658-666.
[29]
LLUGSIR, YACOUBI SEL, FONTAINEA, et al. Comparison between Adam, AdaMax and Adam W optimizers to implement a weather forecast based on neural networks for the andean city of quito[C]//2021 IEEE Fifth Ecuador Technical Chapters Meeting (ETCM). IEEE, 2021: 1-6.
[30]
CUIY, ZENGC, ZHAOX, et al. SportsMOT: A large multi-object tracking dataset in multiple sports scenes[C]//2023 Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). IEEE, 2023: 9887-9897.
[31]
MILANA, LEAL-TAIXÉL, REIDI, et al. MOT16: A benchmark for multi-object tracking[DB/OL].(2016-05-03)[2024-01-27].
[32]
DENDORFERP, REZATOFIGHIH, MILANA, et al. MOT20: A benchmark for multi object tracking in crowded scenes[DB/OL].(2020-03-19)[2024-01-27].