Malicious attackers can easily deceive neural networks by adding human-imperceptible adversarial noise to natural samples, leading to misclassification. To enhance the model’s robustness against such adversarial perturbations, previous research has predominantly concentrated on the robustness of single-modal tasks, with insufficient exploration of multimodal scenarios. Therefore, this paper aims to improve the robustness of multimodal RGB-skeleton action recognition and introduces a robust action recognition framework based on a Feature Interaction Module (FIM), which extracts global information from adversarial samples to learn inter-modal joint representations for calibrating multi-modal features. A corresponding loss function tailored to this framework is also developed. Experimental results demonstrate that against CW attack, our method achieves a RI of 25.14% and an average robust accuracy of 48.99% on the NTURGB+D dataset, outperforming the latest SimMin+ExFMem method by 8.55 and 23.79 percentage points, respectively. These findings confirm that our approach surpasses others in enhancing robustness and balancing accuracy rates.
PALN R, PALS K. A review on image segmentation techniques[J].Pattern Recognition,1993,26(9):1277-1294.
[2]
REDMONJ, DIVVALAS, GIRSHICKR,et al .You only look once:unified,real-time object detection[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). June 27-30,2016,Las Vegas, NV, USA: IEEE, 2016:779-788.
[3]
GLASNERD, BAGONS, IRANIM .Super-resolution from a single image[C]//2009 IEEE 12th International Conference on Computer Vision.September 29-October 2,2009,Kyoto,Japan: IEEE,2009:349-356.
[4]
GOODFELLOWI J, SHLENSJ, SZEGEDYC .Explaining and harnessing adversarial examples[J]. 3rd International Conference on Learning Representations,ICLR 2015-Conference Track Proceedings, 2015: 32-40.
[5]
WEIX X, ZHUJ, YUANS,et al .Sparse adversarial perturbations for videos[J].Proceedings of the AAAI Conference on Artificial Intelligence,2019,33(1):8973-8980.
[6]
WANGH J, WANGG R, LIY,et al .Transferable,controllable,and inconspicuous adversarial attacks on person re-identification with deep mis-ranking[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).June 13-19,2020,Seattle,WA,USA:IEEE,2020:339-348.
[7]
LIL Y, MAR T, GUOQ P,et al .BERT-ATTACK:adversarial attack against BERT using BERT[C]//Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Stroudsburg,PA,USA:Association for Computational Linguistics, 2020: 6193-6202.
[8]
CHENJ Y, YUANB D, TOMIZUKAM .Model-free deep reinforcement learning for urban autonomous driving[C]//2019 IEEE Intelligent Transportation Systems Conference (ITSC).October 27-30,2019,Auckland,New Zealand:IEEE,2019:2765-2771.
[9]
EYKHOLTK, EVTIMOVI, FERNANDESE,et al .Robust physical-world attacks on deep learning visual classification[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition.June 18-23,2018.Salt Lake City,UT,USA:IEEE,2018:1625-1634.
[10]
BUCHV H, AHMEDI, MARUTHAPPUM. Artificial intelligence in medicine:current trends and future possibilities[J]. British Journal of General Practice,2018,68(668):143-144.
[11]
KONGB, WANGX, LIZ Y,et al .Cancer metastasis detection via spatially structured deep network[M]//Lecture Notes in Computer Science.Cham:Springer International Publishing,2017:236-248.
[12]
MAX J, NIUY H, GUL,et al .Understanding adversarial attacks on deep learning based medical image analysis systems[J].Pattern Recognition,2021,110:107332.
[13]
SZEGEDYC, ZAREMBAW, SUTSKEVERI,et al .Intriguing properties of neural networks[EB/OL]. 2013: 1312.6199.
[14]
CARLININ, WAGNERD .Towards evaluating the robustness of neural networks[C]//2017 IEEE Symposium on Security and Privacy (SP).May 22-26,2017,San Jose,CA,USA.IEEE,2017:39-57.
[15]
MADRYA, MAKELOVA, SCHMIDTL, et al. Towards deep learning models resistant to adversarial attacks[C]// 6th International Conference on Learning Representations, ICLR 2018. April 30-May 3, 2018. Vancouver, BC, Canada:OpenReview.net, 2018.
[16]
MADAAND, SHINJ, HWANGS J. Adversarial neural pruning with latent vulnerability suppression[C]//International Conference on Machine Learning, PMLR, 2021: 6575-6585.
[17]
LINJ, GANC, HANS. Defensive quantization: when efficiency meets robustness[C]// 7th International Conference on Learning Representations, ICLR 2019. May 6-9, 2019. New Orleans, LA, USA: OpenReview.net, 2019.
[18]
XIEC H, WUY X, VAN DER MAATENL,et al .Feature denoising for improving adversarial robustness[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).June 15-20,2019,Long Beach,CA,USA:IEEE,2019:501-509.
[19]
NASEERM, KHANS, HAYATM,et al .A self-supervised approach for adversarial robustness[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 13-19,2020,Seattle, WA, USA: IEEE, 2020: 259-268.
[20]
ZHANGJ F, XUX L, HANB,et al .Attacks which do not kill training make adversarial learning stronger[EB/OL].2020:2002.11242.org/abs/2002.11242v2.
[21]
ZHANGJ, ZHUJ, NIUG, et al. Geometry-aware instance-reweighted adversarial training[EB/OL].2021:2010.01736.
BAIY, ZENGY, JIANGY, et al. Improving adversarial robustness via channel-wise activation suppressing[C]// 9th International Conference on Learning Representations, ICLR 2021. May 3-7, 2021. Virtual Event., Austria: OpenReview.net,2021.
[24]
PAPERNOTN, MCDANIELP, WUX,et al .Distillation as a defense to adversarial perturbations against deep neural networks[C]//2016 IEEE Symposium on Security and Privacy (SP). May 22-26,2016,San Jose,CA,USA:IEEE,2016:582-597.
[25]
DAS N, SHANBHOGUEM, CHENS T, et al. Compression to the rescue: defending from adversarial attacks across modalities[C]//ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 2018.
[26]
WANGL, HEZ, TANGJ, et al. A dual semantic-aware recurrent global-adaptive network for vision-and-language navigation[C]// Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence. August 19-25,2023.Macao: International Joint Conferences on Artificial Intelligence Organization, 2023: 1479-1487.
[27]
FUZ, MAOZ, SONGY, et al. Learning semantic relationship among instances for image-text matching[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).June 17-24,2023,Vancouver,BC,Canada: IEEE,2023:15159-15168.
[28]
KHADERF, MUELLER-FRANZESG, WANGT, et al. Medical diagnosis with large scale multimodal transformers: leveraging diverse data for more accurate diagnosis [EB/OL].2022: 2212.09162.https: arxiv.org/abs/2212.09162.
[29]
MOONJ H, LEEH, SHINW,et al .Multi-modal understanding and generation for medical images and text via vision-language pre-training[J].IEEE Journal of Biomedical and Health Informatics,2022,26(12):6070-6080.
ILYASA, SANTURKARS, TSIPRASD, et al. Adversarial examples are not bugs, they are features[C]// Proceedings of the 33rd International Conference on Neural Information Processing Systems. Dec,2020. ACM,2020: 125-136.
[32]
KINFUK A, VIDALR .Analysis and extensions of adversarial training for video classification[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). June 19-20,2022,New Orleans,LA,USA.IEEE,2022:3415-3424.
[33]
ZHANGJ M, YIQ, SANGJ T .Towards adversarial attack on vision-language pre-training models[C]//Proceedings of the 30th ACM International Conference on Multimedia. Lisboa Portugal:ACM, 2022: 5005-5013.
[34]
ZHAOY Q, PANGT Y, DUC,et al .On evaluating adversarial robustness of large vision-language models[EB/OL]. 2023:2305.16934.org/abs/2305.16934v2.
[35]
DUANH D, ZHAOY, CHENK,et al .Revisiting skeleton-based action recognition[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).June 18-24,2022,New Orleans, LA, USA: IEEE, 2022: 2959-2968.
[36]
YUB X B, LIUY, ZHANGX,et al .MMNet:a model-based multimodal network for human action recognition in RGB-D videos[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence,2023,45(3):3522-3538.
[37]
VAEZI JOZEH R, SHABANA, IUZZOLINOM L,et al .MMTM:multimodal transfer module for CNN fusion[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).June 13-19,2020,Seattle,WA,USA: IEEE,2020:13286-13296.
[38]
TIANY P, XUC L .Can audio-visual integration strengthen robustness under multimodal attacks?[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).June 20-25,2021,Nashville,TN,USA: IEEE,2021:5597-5607.
[39]
YANS J, XIONGY J, LIND H .Spatial temporal graph convolutional networks for skeleton-based action recognition[J].Proceedings of the AAAI Conference on Artificial Intelligence,2018,32(1):9-18.
[40]
CARREIRAJ, ZISSERMANA .Quo vadis,action recognition?A new model and the kinetics dataset[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).July 21-26,2017,Honolulu,HI,USA: IEEE, 2017: 4724-4733.
[41]
LIC, ZHONGQ Y, XIED,et al .Co-occurrence feature learning from skeleton data for action recognition and detection with hierarchical aggregation[C]//Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence.July 13-19,2018.Stockholm,Sweden.California:International Joint Conferences on Artificial Intelligence Organization,2018:786-792.
[42]
SHAHROUDYA, LIUJ, NGT T,et al .NTU RGB D:a large scale dataset for 3D human activity analysis[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).June 27-30,2016,Las Vegas,NV,USA: IEEE,2016:1010-1019.
[43]
LIUX, SHIH L, CHENH Y,et al .iMiGUE:an identity-free video dataset for micro-gesture understanding and emotion analysis[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).June 20-25,2021,Nashville,TN,USA: IEEE, 2021: 10626-10637.
[44]
HUJ, SHENL, SUNG .Squeeze-and-excitation networks[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition.June 18-23,2018,Salt Lake City,UT,USA: IEEE,2018: 7132-7141.
基金资助
国家自然科学基金资助项目(62072331)
National Natural ScienceFoundation of China(62072331)
国家自然科学基金资助项目(62231018)
National Natural Science Foundation of China(62231018)
国家自然科学基金资助项目(62171309)
National Natural Science Foundation of China(62171309)