IIG-VSAD:基于实例信息引导的视频流行为检测方法

夏子瀛 ,  头旦才让 ,  张艺杰 ,  刘思宇 ,  邓昌健 ,  程建 ,  尼玛扎西

电子科技大学学报 ›› 2026, Vol. 55 ›› Issue (1) : 129 -136.

PDF (1367KB)
电子科技大学学报 ›› 2026, Vol. 55 ›› Issue (1) : 129 -136. DOI: 10.12178/1001-0548.2024328
计算机工程与应用

IIG-VSAD:基于实例信息引导的视频流行为检测方法

作者信息 +

IIG-VSAD: Instance information-guided video stream action detection

Author information +
文章历史 +
PDF (1399K)

摘要

视频流时序行为检测要求在仅观测到历史及当前时空信息的条件下,于在线视频流中准确检测当前时刻的行为类别。现有方法主要通过设计网络并利用帧级信息进行监督学习,对单帧信息过度敏感,缺乏时序一致性,导致检测准确性不足。针对以上问题,提出实例信息引导的视频流行为检测方法,在单帧检测基础上扩增实例信息,提出实例图推理策略生成导引,融合时序特征提升检测性能。于公开视频数据验证实验结果,证明了该方法的有效性且具备高效的检测效率。

Abstract

Action detection in video streams requires accurately identifying the action category at the current moment within an online video stream, given only the historical and current spatiotemporal information observed up to that point. The existing methods mainly conduct supervised learning by designing networks and using frame-level information, which are overly sensitive to single-frame information and lack temporal consistency, resulting in insufficient detection accuracy. To address the aforementioned issues, an instance-guided video stream action detection method is proposed. Building upon frame-level detection, instance information is augmented, an instance graph reasoning strategy is proposed for generating guidance, and temporal features are then integrated to enhance detection performance. Finally, the proposed algorithm is validated on publicly available video datasets, and experimental results demonstrate the effectiveness of the method and its high detection efficiency.

关键词

视频流 / 行为检测 / 实例导引生成 / 实例图推理 / 注意力机制

Key words

video stream / action detection / instance guidance generation / instance graph reasoning / attention mechanism

引用本文

引用格式 ▾
夏子瀛,头旦才让,张艺杰,刘思宇,邓昌健,程建,尼玛扎西. IIG-VSAD:基于实例信息引导的视频流行为检测方法[J]. 电子科技大学学报, 2026, 55(1): 129-136 DOI:10.12178/1001-0548.2024328

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

SHU T M, XIE D, ROTHROCK B, et al. Joint inference of groups, events and human roles in aerial videos[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. [S.l.]. IEEE,2015: 4576-4584.

[2]

翟强, 程洪, 黄瑞, . 智能汽车中人工智能算法应用及其安全综述[J].电子科技大学学报, 2020, 49(4): 490-498

[3]

ZHAI Q, CHENG H, HUANG R, et al. A survey: Artificial intelligence and its security in intelligent vehicle[J].Journal of University of Electronic Science and Technology of China, 2020, 49(4): 490-498.

[4]

MU Y, ZHANG Q, HU M, et al. Embodiedgpt: Vision—language pre—training via embodied chain of thought[EB/OL]. [2024—09—23].https://openreview.net/pdf?id=IL5zJqfxAa.

[5]

王柏村, 薛塬, 延建林, . 以人为本的智能制造: 理念、技术与应用[J].中国工程科学, 2020, 22(4): 139-146

[6]

WANG B C, XUE Y, YAN J L, et al. Human—centered intelligent manufacturing: Overview and perspectives[J].Strategic Study of CAE, 2020, 22(4): 139-146.

[7]

SHI D F, ZHONG Y J, CAO Q, et al. ReAct: Temporal action detection with relational queries[M]//Computer Vision — ECCV 2022. Cham: Springer, 2022: 105-121.

[8]

CARION N, MASSA F, SYNNAEVE G, et al. End—to—end object detection with transformers[M]//Computer Vision — ECCV 2020. Cham: Springer, 2020: 213-229.

[9]

ZHU Z, WANG L, TANG W, et al. ContextLoc++: A unified context model for temporal action localization[J].IEEE Trans Pattern Anal Mach Intell, 2023, 45(8): 9504-9519.

[10]

LIN T W, LIU X, LI X, et al. BMN: Boundary—matching network for temporal action proposal generation[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. [S.l.]: IEEE,2019: 3889-3898.

[11]

FOO L G, LI T J, RAHMANI H, et al. Action detection via an image diffusion process[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. [S.l.]: IEEE,2024: 18351-18361.

[12]

LIU S M, ZHANG C L, ZHAO C, et al. End—to—end temporal action detection with 1B parameters across 1000 frames[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. [S.l.]: IEEE,2024: 18591-18601.

[13]

VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[J].Advances in Neural Information Processing Systems, 2017, 30: 5998-6008.

[14]

XU M Z, GAO M F, CHEN Y T, et al. Temporal recurrent networks for online action detection[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. [S.l.]: IEEE,2019: 5532-5541.

[15]

AN J, KANG H, HAN S H, et al. MiniROAD: Minimal RNN framework for online action detection[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. [S.l.]: IEEE,2023: 10341-10350.

[16]

WANG X, ZHANG S W, QING Z W, et al. OadTR: Online action detection with transformers[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. [S.l.]: IEEE,2021: 7565-7575.

[17]

LIU T S, LAM K M, BAO B K. A memory—assisted knowledge transferring framework with curriculum anticipation for weakly supervised online activity detection[J].International Journal of Computer Vision, 2025, 133(4): 1940-1963.

[18]

WANG L M, XIONG Y J, WANG Z, et al. Temporal segment networks: Towards good practices for deep action recognition[M]//Computer Vision — ECCV 2016. Cham: Springer International Publishing, 2016: 20-36.

[19]

ZHENG Z H, WANG P, LIU W, et al. Distance—IoU loss: Faster and better learning for bounding box regression[J].Proceedings of the AAAI Conference on Artificial Intelligence, 2020, 34(7): 12993-13000.

[20]

BODLA N, SINGH B, CHELLAPPA R, et al. Soft—NMS: Improving object detection with one line of code[C]//Proceedings of the IEEE International Conference on Computer Vision. [S.l.]: IEEE,2017: 5561-5569.

[21]

IDREES H, ZAMIR A R, JIANG Y G, et al. The THUMOS challenge on action recognition for videos “in the wild”[J].Computer Vision and Image Understanding, 2017, 155: 1-23.

[22]

CARREIRA J, ZISSERMAN A. Quo vadis, action recognition? A new model and the kinetics dataset[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. [S.l.]: IEEE,2017: 6299-6308.

[23]

HE K M, ZHANG X Y, REN S Q, et al. Deep residual learning for image recognition[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. [S.l.]: IEEE,2016: 770-778.

[24]

IOFFE S, SZEGEDY C. Batch normalization: Accelerating deep network training by reducing internal covariate shift[C]//International Conference on Machine Learning. [S.l.]: [s.n.],2015: 448-456.

[25]

MIN S, MOON J. Information elevation network for online action detection and anticipation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. [S.l.]: IEEE,2022: 2550-2558.

[26]

YANG L, HAN J W, ZHANG D W. Colar: Effective and efficient online action detection by consulting exemplars[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. [S.l.]: IEEE,2022: 3160-3169.

[27]

ZHU Z Q, SHAO W C, JIAO D D. TLS—RWKV: Real—time online action detection with temporal label smoothing[J].Neural Processing Letters, 2024, 56(2): 57.

基金资助

国家自然科学基金民航联合基金重点项目(U2233209)

国家自然科学基金青年项目(62306158)

四川省自然科学基金(2023NSFSC0484)

AI Summary AI Mindmap
PDF (1367KB)

209

访问

0

被引

详细

导航
相关文章

AI思维导图

/