多模态大模型在专业体育运动中的研究

陈旭, 郭羽翔, 杨艳丽, 郭浩

小型微型计算机系统 ›› 2026, Vol. 47 ›› Issue (9) : 2158 -2164.

小型微型计算机系统 ›› 2026, Vol. 47 ›› Issue (9) : 2158 -2164. DOI: 10.20009/j.cnki.21-1106/TP.2025-0299
算法理论与人工智能

多模态大模型在专业体育运动中的研究

    陈旭1, 郭羽翔2, 杨艳丽1, 郭浩1
作者信息 +

Research on Multimodal Large Models in Professional Sports

    CHEN Xu1, GUO Yuxiang2, YANG Yanli1, GUO Hao1
Author information +
文章历史 +

摘要

随着体育视频分析应用的拓展,传统计算机视觉方法在高速运动、复杂光照变化及图像模糊等场景中表现出鲁棒性不足,难以满足对精确事件识别与智能理解的需求.针对上述问题,研究提出了一种基于多模态大模型的体育场景理解方法ERS-MLLM(Event RGB Sports-Multimodal Large Language Model),通过处理事件相机数据与传统RGB视频数据,实现多源信息的深度整合.该模型利用跨模态学习机制,在保持时间动态信息的同时增强了空间语义的表达能力,从而提高了体育场景中事件识别、行为预测与决策推理的准确性与稳定性.实验在SPORTM数据集上开展性能评估,结果显示ERS-MLLM在事件识别精度、行为预测准确率与跨模态推理能力方面均优于现有主流多模态模型,尤其在高速物体运动等复杂场景下表现更为出色.

Abstract

As the application of sports video analysis expands,traditional computer vision methods show limited robustness in scenarios involving high-speed motion,complex lighting variations,and image blur.These limitations hinder accurate event recognition and intelligent understanding.To address these challenges,this paper proposes a sports scene understanding method based on a multimodal large language model,named ERS-MLLM (Event RGB Sports-Multimodal Large Language Model).The method integrates data from event cameras and conventional RGB videos to achieve deep fusion of multi-source information.ERS-MLLM applies a cross-modal learning mechanism to enhance spatial semantic representation while preserving temporal dynamics.This design improves the accuracy and stability of event recognition,behavior prediction,and decision reasoning in sports scenarios.The performance of the experiment was evaluated on the SPORTM dataset.The results showed that ERS-MLLM outperformed the existing mainstream multimodal models in terms of event recognition accuracy,behavior prediction accuracy,and cross-modal reasoning ability,especially in complex scenarios such as high-speed object movement.

关键词

多模态大模型 / 事件相机 / 体育视频分析 / 特征对齐 / 智能推理

Key words

multimodal large language model / event camera / sports video analysis / feature alignment / intelligent reasoning

引用本文

引用格式 ▾
陈旭, 郭羽翔, 杨艳丽, 郭浩. 多模态大模型在专业体育运动中的研究[J]. 小型微型计算机系统, 2026, 47(9): 2158-2164 DOI:10.20009/j.cnki.21-1106/TP.2025-0299

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1] He Y C,Yuan Z Q,Wu Y H,et al.Vistec:video modeling for sports technique recognition and tactical analysis [C]//Proceedings of the AAAI Conference on Artificial Intelligence,2024:8490-8498.
[2] Xia H T,Tracy R,Zhao Y,et al.Advanced volleyball stats for all levels:automatic setting tactic detection and classification with a single camera [C]//IEEE International Conference on Data Mining Workshops,2023:1407-1416.
[3] Gehrig D,Scaramuzza D.Low-latency automotive vision with event cameras[J].Nature,2024,629(8014):1034-1040.
[4] Gehrig D,Rebecq H,Gallego G,et al.EKLT:asynchronous photometric feature tracking using events and frames[J].International Journal of Computer Vision,2020,128(3):601-618.
[5] LIN E X,ZHANG C,ZHOU X T,et al.Research on large industrial vision models based on ensemble self-supervised learning[J].Journal of Chinese Computer Systems,2025,46(4):907-913.
[6] Li J N,Li D X,Xiong C M,et al.BLIP:bootstrapping language-image pre-training for unified vision-language understanding and generation[C]//Proceedings of the International Conference on Machine Learning,2022:12888-12900.
[7] Alayrac J B,Donahue J,Luc P,et al.Flamingo:a visual language model for few-shot learning[C]//Proceedings of the 35th Conference on Neural Information Processing Systems,2022:23716-23736.
[8] Chen Z,Wang W,Tian H,et al.How far are we to GPT-4V? Closing the gap to commercial multimodal models with open-source suites[J].Science China Information Sciences,2024,67(12):220101,doi:10.1007/s11432-024-4231-5.
[9] Huang K H,Li C,Chang K W.Generating sports news from live commentary:a Chinese dataset for sports game summarization [C]//Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing,2020:609-615.
[10] Vujicic S,Mladenovic M.An approach to automatic classification of hate speech in sports domain on social media [J].Journal of Big Data,2023,10(1):1-16.
[11] Held J,Cioppa A,Giancola S,et al.VARS:video assistant referee system for automated soccer decision making from multiple views[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2023:5086-5097.
[12] Jardim P C,Moraes L M P,Aguiar C D.QaSports:a question answering dataset about sports[C]//Anais do V Dataset Showcase Workshop,2023:1-12.
[13] Held J,Itani H,Cioppa A,et al.X-VARS:introducing explainability in football refereeing with multi-modal large language models[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2024:3267-3279.
[14] Cho H,Kim H,Chae Y,et al.Label-free event-based object recognition via joint learning with image reconstruction from events[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision,2023:19866-19877.
[15] Chen Y Y,Sikka K,Cogswell M,et al.DRESS:instructing large vision-language models to align and interact with humans via natural language feedback[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2024:14239-14250.
[16] ZHANG Y W,ZHOU Q,CHEN W,et al.Entity alignment on multi-modal knowledge graph [J].Journal of Chinese Computer Systems,2024,45(5):1257-1263.
[17] Chen Z,Wu J N,Wang W H,et al.InternVL:scaling up vision foundation models and aligning for generic visual-linguistic tasks[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2024:24185-24198.
[18] Zhang Y W,Zhou Q,Chen W,et al.Entity alignment on multi-modal knowledge graph [J].Journal of Chinese Computer Systems,2024,45(5):1257-1263.
[19] Singh A,Hu R H,Goswami V,et al.FLAVA:a foundational language and vision alignment model[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2022:15638-15650.
[20] Yu B H,Ren J J,Han J,et al.EventPS:real-time photometric stereo using an event camera[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,2024:9602-9611.
附中文参考文献:
[5] 林而贤,张 潮,周雄图,等.基于集成自监督的工业视觉大模型算法研究[J].小型微型计算机系统,2025,46(4):907-913.
[16] 张艺玮,周 乾,陈 伟,等.面向多模态知识图谱的实体对齐方法研究[J].小型微型计算机系统,2024,45(5):1257-1263.

基金资助

国家自然科学基金项目(62403345)资助.

AI Summary AI Mindmap

0

访问

0

被引

详细

导航
相关文章

AI思维导图

/