To comprehensively assess the robustness of Tibetan language models, this study introduces a Tibetan adversarial sample generation method named DAAttacker. This approach employs a multi-granularity perturbation strategy at both the character and syllable levels, incorporating six types of perturbations, including insertion, deletion, and substitution, to generate diverse adversarial samples. Under a white-box attack setting, DAAttacker accurately identifies critical Tibetan syllables for perturbation and utilizes a scoring mechanism to obtain optimal perturbations, while a similarity threshold is applied to maintain semantic fidelity and ensure balance between naturalness and stealthiness. Experimental results on multiple Tibetan classification models show that DAAttacker achieves attack success rates exceeding 90%, and the generated adversarial samples exhibit notable fluency and concealment. These findings provide new insights into adversarial sample generation and model robustness enhancement for low-resource languages.
在对抗样本生成过程中,扰动对分类器预测结果的影响尤为重要。Ren等[21]提出了基于预测概率变化(Probability Weighted Word Saliency,PWWS)的词重要性排序策略,证明了该指标能有效筛选出最具攻击性的词汇。因此,通过计算扰动前后的预测概率变化ΔP来进一步评估扰动效果。本文通过加权Jaccard相似度和预测概率变化的综合得分来评估每个扰动的效果。综合得分指标S可以表示为:
KrizhevskyA, SutskeverI, HintonG E. ImageNet classification with deep convolutional neural networks[J]. Communications of the ACM, 2017, 60(6):84-90.
[2]
BowmanS, VilnisL, VinyalsO, et al. Generating sentences from a continuous space[C]. Proceedings of the 20th SIGNLL conference on computational natural language learning, 2016: 10-21..
[3]
MilneD, WittenI H. An open-source toolkit for mining Wikipedia[J]. Artificial Intelligence, 2013, 194: 222-239.
[4]
EhsanU, HarrisonB, ChanL, et al. Rationalization: A neural machine translation approach to generating natural language explanations[C]. Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, 2018: 81-87.
[5]
IyerN, PradhanA. " Was there Bias in your response?": LLM anthropomorphic traits in context of ability Bias questioning[C]. Companion Publication of the 2025 Conference on Computer-Supported Cooperative Work and Social Computing, 2025: 471-476.
[6]
LingL, RabbiF, WangS, et al. Bias revealed: Investigating social bias in LLM-Generated Code[C]. Proceedings of the AAAI Conference on Artificial Intelligence. 2025, 39(26): 27491-27499.
[7]
JiaR, LiangP. Adversarial examples for evaluating reading comprehension systems[C]. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, 2017: 2021-2031.
[8]
EbrahimiJ, RaoA, LowdD, et al. Hotflip: White-box adversarial examples for text classification[C]. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2018: 31-36.
[9]
GaoJ, LanchantinJ, SoffaM L, et al. Black-box generation of adversarial text sequences to evade deep learning classifiers[C]. 2018 IEEE Security and Privacy Workshops (SPW) IEEE, 2018: 50-56.
[10]
WallaceE, FengS, KandpalN, et al. Universal adversarial triggers for attacking and analyzing NLP[C]. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-JCNLP), 2019:2153-2162.
[11]
LiJ, JiS, DuT, et al. Textbugger: Generating adversarial text against real-world applications[J]. arXiv preprint arXiv:2018.
[12]
TongX, WangL, WangR, et al. A generation method of word-level adversarial samples for Chinese text classification[J]. Netinfo Secur, 2020, 20(9): 12-16.
[13]
HanZ Y, WangW, XuanS C. Chinese adversarial example generation guided by multi-constraints[J]. Journal of Chinese Information Processing, 2023, 37(2): 41-52.
[14]
CaoX, DawaD, QunN, et al. Pay attention to the robustness of Chinese minority language models! Syllable-level textual adversarial attack on Tibetan script[C]. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023), 2023: 35-46.
[15]
WangC, ZengJ, WuC. Generating fluent Chinese adversarial examples for sentiment classification[C]. 2020 IEEE 14th International Conference on Anti-counterfeiting, Security, and Identification (ASID). IEEE, 2020: 149-154.
[16]
ZhangJ, KazhuoD, GadengL, et al. Research and application of tibetan pre-training language model based on bert[C]. Proceedings of the 2022 2nd International Conference on Control and Intelligent Robotics, 2022: 519-524.
[17]
YangZ, XuZ, CuiY, et al. CINO: A Chinese minority pre-trained language model[J]. arXiv preprint arXiv:2022.
[18]
PapernotN, McDanielP, JhaS, et al. The limitations of deep learning in adversarial settings[C]. 2016 IEEE European Symposium on Security and Rrivacy (EuroSP). IEEE, 2016: 372-387.
[19]
GhorbaniH. Mahalanobis distance and its application for detecting multivariate outliers[J]. Facta Universitatis, Series: Mathematics and Informatics, 2019: 583-595.
[20]
FletcherS, IslamM Z. Comparing sets of patterns with the Jaccard index[J]. Australasian Journal of Information Systems, 2018, 22.
[21]
RenS, DengY, HeK, et al. Generating natural language adversarial examples through probability weighted word saliency[C]. Proceedings of the 57th annual meeting of the association for computational linguistics, 2019: 1085-1097.
[22]
GraveE, BojanowskiP, GuptaP, et al. Learning word vectors for 157 languages[C]. Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018),2018:1690-1696.
[23]
CerD, YangY, KongS, et al. Universal sentence encoder for English[C]. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2018: 169-174.
[24]
QunN, LiX, QiuX, et al. End-to-end neural text classification for tibetan[C]. International Symposium on Natural Language Processing Based on Naturally Annotated Big Data. Cham: Springer International Publishing, 2017: 472-480.
[25]
ZhuY, DejiK, QunN, et al. Sentiment analysis of tibetan short texts based on graphical neural networks and pre-training models[J]. Journal of Chinese Information Processing, 2023, 37(2): 71-79.