Key Laboratory of Aerospace Information Security and Trusted Computing,Ministry of Education,School of Cyber Science and Engineering,Wuhan University,Wuhan 430072,Hubei,China
Existing decision-based black-box text adversarial attack methods cannot balance attack effectiveness and efficiency. Therefore, a simple and efficient decision-based word-level black-box text adversarial attack method called TextLeak is proposed. The fundamental concept of this method involes searching for the minimum perturbation required to generate adversarial examples through a multi-level search. Specifically, it begins with a coarse-grained search to identify the target area. Then, it uses fine-grained search based on this target area to find the optimal solution as the adversarial example. The main evaluation metrics are attack success rate, perturbed rate, and query number. Three state-of-the-art decision-based black-box text adversarial attacks are selected as baseline methods for experimental comparison on the same dataset and model. The experimental results show that TextLeak has an average query number of about 368 times and an average attack success rate of about 96.0% on text classification tasks. Compared with the population-based optimization algorithm(POA), TextLeak has an average query number of about 5.25% of POA while maintaining a comparable attack success rate. This demonstrates that TextLeak has a high attack success rate and query efficiency, and it is a simple, efficient, and practical text adversarial attack method with broad application prospects.
KRIZHEVSKYA, SUTSKEVERI, HINTONG E. ImageNet classification with deep convolutional neural networks[J].Communications of the ACM, 2017, 60(6): 84-90. DOI:10.1145/3065386 .
[2]
GRAVESA, MOHAMEDA R, HINTONG. Speech recognition with deep recurrent neural networks[C]//2013 IEEE International Conference on Acoustics, Speech and Signal Processing. New York: IEEE Press, 2013: 6645-6649. DOI: 10.1109/ICASSP.2013.6638947 .
[3]
MIKOLOVT, KARAFIÁTM, BURGETL, et al. Recurrent neural network based language model[EB/OL]. [2021-09-23]. DOI: 10.21437/interspeech.2010-343 .
[4]
SZEGEDYC, ZAREMBAW, SUTSKEVERI, et al. Intriguing properties of neural networks[EB/OL]. [2021-10-16].
[5]
PAPERNOTN, MCDANIELP, SWAMIA, et al. Crafting adversarial input sequences for recurrent neural networks[C]// MILCOM 2016-2016 IEEE Military Communications Conference. New York: IEEE Press, 2016: 49-54. DOI: 10.1109/MILCOM.2016.7795300 .
[6]
EBRAHIMIJ, RAOA Y, LOWDD, et al. Hotflip: White-box adversarial examples for text classification[C]//Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics(Volume 2: Short Papers). Stroudsburg: Association for Computational Linguistics, 2018: 31-36. DOI: 10.18653/v1/p18-2006 .
[7]
WANGB X, XUC J, LIUX Y, et al. SemAttack: Natural textual attacks via different semantic spaces[C]//Findings of the Association for Computational Linguistics: NAACL 2022. Stroudsburg: Association for Computational Linguistics, 2022: 176-205. DOI: 10.18653/v1/2022.findings-naacl.14 .
[8]
ALZANTOTM, SHARMAY, ELGOHARYA, et al. Generating natural language adversarial examples[C]// Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Stroudsburg: Association for Computational Linguistics, 2018: 2890-2896. DOI: 10.18653/v1/d18-1316 .
[9]
JIND, JINZ J, ZHOUJ T, et al. Is BERT really robust? A strong baseline for natural language attack on text classification and entailment[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2020, 34(5): 8018–8025. DOI: 10.1609/aaai.v34i05.6311 .
[10]
RENS H, DENGY H, HEK, et al. Generating natural language adversarial examples through probability weighted word saliency[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Stroudsburg: Association for Computational Linguistics, 2019: 1085-1097. DOI: 10.18653/v1/p19-1103 .
[11]
LIJ F, JIS L, DUT Y, et al. TextBugger: Generating adversarial text against real-world applications[C]//Proceedings 2019 Network and Distributed System Security Symposium. Reston: Internet Society, 2019: 1-15. DOI: 10.14722/ndss.2019.23138 .
[12]
LEED, MOONS, LEEJ, et al. Query-efficient and scalable black-box adversarial attacks on discrete sequential data via Bayesian optimization[EB/OL]. 2022: arXiv: 2206.08575.
MAHESHWARYR, MAHESHWARYS, PUDIV. Generating natural language attacks in a hard label black box setting[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2021, 35(15):13525-13533. DOI: 10.1609/aaai.v35i15.17595 .
[15]
SAXENAS. TextDecepter: Hard label black box attack on text classifiers[EB/OL]. 2020: arXiv: 2008.06860.
[16]
YEM C, CHENJ H, MIAOC L, et al. LeapAttack: Hard-label adversarial attack on text via gradient-based optimization[C]//Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. New York: ACM, 2022: 2307-2315. DOI: 10.1145/3534678.3539357 .
[17]
MRKŠIĆN, SÉAGHDHAD Ó, THOMSONB, et al. Counter-fitting word vectors to linguistic constraints[C]//Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Stroudsburg: Association for Computational Linguistics, 2016: 142-148. DOI: 10.18653/v1/n16-1018 .
[18]
ZHANGX, ZHAOJ, LECUNY. Character-level convolutional networks for text classification[EB/OL]. [2020-06-15]. DOI: 10.48550/arXiv.1509.01626 .
[19]
MAASA L, DALYR E, PHAMP T, et al. Learning word vectors for sentiment analysis [C]//Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies - Volume 1. New York: ACM, 2011: 142–150. DOI: 10.5555/2002472.2002491 .
[20]
PANGB, LEEL. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales [C]//Proceedings of the 43rd Annual Meeting on Association for Computational Linguistics. New York: ACM, 2005:115-124. DOI: 10.3115/1219840.1219855 .
[21]
KIMY. Convolutional neural networks for sentence classification[C]//Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). Stroudsburg: Association for Computational Linguistics, 2014: 1746–1751. DOI: 10.3115/v1/d14-1181 .
DEVLINJ, CHANGM, LEEK, et al. Bert: Pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics. Stroudsburg: Association for Computational Linguistics, 2019: 4171-4186. DOI: 10.18653/v1/N19-1423 .
[24]
MIYATOT, DAIA M, IAN G. Adversarial training methods for semi-supervised text classification[EB/OL]. [2022-09-13].
[25]
MADRYA, MAKELOVA, SCHMIDTL, et al. Towards deep learning models resistant to adversarial attacks[EB/OL]. [2019-03-24]. DOI: 10.48550/arXiv.1706.06083 .