Existing studies on backdoor detection mainly focus on white-box settings. However, internal information is not always exposed to the security auditors. To bridge this gap, we investigated the black-box backdoor detection based on Monte-Carlo gradient estimation. We propose a black-box trigger reversing method to determine whether a model is infected with a backdoor attack by regarding the black-box trigger reversing problem as a zeroth-order optimization. Furthermore, we propose a fast black-box backdoor detection method to reduce the detection overhead by using the importance sampling technique, a regularized criterion, and an early stopping strategy. Experimental results on three popular image datasets show that the proposed method can accurately distinguish benign and infected models and obtain effective backdoor triggers.
采用检测准确率和AUROC(the area under receiver operating curve)常用于分类任务中的指标来评估最终的检测结果。检测准确率只考虑检测正确的样本占所选样本集合的比例,而AUROC综合考虑正类样本和负类样本被区分开的程度,最理想的情况下这两个指标都达到最大值1.0。表2展示了最终的检测结果,白盒方法NC在3个数据集上的检测准确率达到了100%,AUROC值达到1.0,而黑盒触发器逆向方法在MNIST和GTSRB数据集上的检测准确率达到了100%,AUROC值达到1.0,在CIFAR-10数据集上出现了少量错误,检测准确率和AUROC值分别为99.5%和0.996。这个结果说明基于黑盒触发器逆向的后门检测方法具有非常好的分类性能,接近于白盒方法,达成了判断模型是否被植入后门的防御目标。
WANGM, DENGW H. Deep face recognition: A survey [J]. Neurocomputing, 2021, 429: 215-244. DOI:10.1016/j.neucom.2020.10.081 .
[2]
LIUL, OUYANGW L, WANGX G, et al. Deep learning for generic object detection: A survey [J]. International Journal of Computer Vision, 2020, 128(2): 261-318. DOI:10.1007/s11263-019-01247-4 .
[3]
GRIGORESCUS, TRASNEAB, COCIAST, et al. A survey of deep learning techniques for autonomous driving[J]. Journal of Field Robotics, 2020, 37(3): 362-386. DOI:10.1002/rob.21918 .
[4]
SHEND G, WUG R, SUKH I. Deep learning in medical image analysis [J]. Annual Review of Biomedical Engineering, 2017, 19: 221-248. DOI:10.1146/annurev-bioeng-071516-044442 .
[5]
HASKINSG, KRUGERU, YANP K. Deep learning in medical image registration: A survey [J]. Machine Vision and Applications, 2020, 31(1/2): 1-18. DOI:10.1007/s00138-020-01060-x .
[6]
GUT Y, LIUK, DOLAN-GAVITTB, et al. BadNets: Evaluating backdooring attacks on deep neural networks [J]. IEEE Access, 2019, 7: 47230-47244. DOI:10.1109/ACCESS.2019.2909068 .
[7]
LIUY Q, MAS Q, AAFERY, et al. Trojaning attack on neural networks [C]//Proceedings 2018 Network and Distributed System Security Symposium. Reston: Internet Society, 2018:3A5. DOI:10.14722/ndss.2018.23291 .
[8]
SAHAA, SUBRAMANYAA, PIRSIAVASHH. Hidden trigger backdoor attacks [J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2020, 34(7): 11957-11965. DOI:10.1609/aaai.v34i07.6871 .
YAOY S, LIH Y, ZHENGH T, et al. Latent backdoor attacks on deep neural networks[C]//Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security. New York: ACM, 2019: 2041-2055. DOI:10.1145/3319535.3354209 .
[11]
CHOUE, TRAMÈRF, PELLEGRINOG. SentiNet: detecting localized universal attacks against deep learning systems [C]//2020 IEEE Security and Privacy Workshops. New York: IEEE Press, 2020: 48-54. DOI:10.1109/SPW50608.2020.00025 .
[12]
GAOY S, XUC G, WANGD R, et al. STRIP: A defence against Trojan attacks on deep neural networks[C]//Proceedings of the 35th Annual Computer Security Applications Conference. New York: ACM, 2019: 113-125. DOI:10.1145/3359789.3359790 .
[13]
TRANB, LIJ, MADRYA. Spectral signatures in backdoor attacks [C]//Proceedings of Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018. Montréal: MIT Press. 2018: 8011-8021.
[14]
YAOY S, LIH Y, ZHENGH T,et al. Latent backdoor attacks on deep neural networks [EB/OL].[2019-01-20]. DOI: 10.1145/3319535.3354209 .
[15]
WANGB L, YAOY S, SHANS, et al. Neural cleanse: Identifying and mitigating backdoor attacks in neural networks[C]//2019 IEEE Symposium on Security and Privacy. New York: IEEE Press, 2019: 707-723. DOI:10.1109/SP.2019.00031 .
[16]
CHENH L, FUC, ZHAOJ S, et al. DeepInspect: A black-box Trojan detection and mitigation framework for deep neural networks [C]//Proceedings of the 28th International Joint Conference on Artificial Intelligence. New York: ACM, 2019: 4658-4664. DOI:10.24963/ijcai.2019/647 .
[17]
LIUY Q, LEEW C, TAOG H, et al. ABS: scanning neural networks for back-doors by artificial brain stimulation [C]//Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security. New York: ACM, 2019: 1265-1282. DOI:10.1145/3319535.3363216 .
[18]
XUX J, WANGQ, LIH C, et al. Detecting AI Trojans using meta neural analysis[C]//2021 IEEE Symposium on Security and Privacy. New York: IEEE Press, 2021: 103-120. DOI:10.1109/SP40001.2021.00034 .
[19]
KOLOURIS, SAHAA, PIRSIAVASHH, et al. Universal litmus patterns: Revealing backdoor attacks in CNNs [C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2020: 298-307. DOI:10.1109/CVPR42600.2020.00038 .
[20]
SZEGEDYC, ZAREMBAW, SUTSKEVERI, et al. Intriguing Properties of Neural Networks [DB/OL]. [2021-02-20].
[21]
KINGMAD P, BAJ L. Adam: A Method for Stochastic Optimization [DB/OL].[2021-02-20].
GER, HUANGF R, JINC, et al. Escaping from Saddle Points — Online Stochastic Gradient for Tensor Decomposition [EB/OL]. 2015: arXiv: 1503.02101.
[24]
GHADIMIS, LANG H. Stochastic first- and zeroth-order methods for nonconvex stochastic programming [J]. SIAM Journal on Optimization, 2013, 23(4): 2341-2368. DOI: 10.1137/120880811 .
[25]
NESTEROVY, SPOKOINYV. Random gradient-free minimization of convex functions [J]. Foundations of Computational Mathematics, 2017, 17(2): 527-566. DOI:10.1007/s10208-015-9296-2 .
[26]
CHENP Y, ZHANGH, SHARMAY, et al. ZOO: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models [C]// Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security. New York:Association for Computing Machinery, 2017: 15-26. DOI:10.1145/3128572.3140448 .
[27]
CHENJ B, JORDANM I, WAINWRIGHTM J. HopSkipJumpAttack: A query-efficient decision-based attack[C]//2020 IEEE Symposium on Security and Privacy. New York: IEEE Press, 2020: 1277-1294. DOI:10.1109/SP40000.2020.00045 .
[28]
LECUNY, BOTTOUL, BENGIOY, et al. Gradient-based learning applied to document recognition [J]. Proceedings of the IEEE, 1998, 86(11): 2278-2324. DOI:10.1109/5.726791 .
[29]
KRIZHEVSKYA. Learning Multiple Layers of Features from Tiny Images [EB/OL].[2022-04-01]. DOI: 10.1016/j.tics.2007.09.004 .
[30]
HOUBENS, STALLKAMPJ, SALMENJ, et al. Detection of traffic signs in real-world images: The German traffic sign detection benchmark[C]//The 2013 International Joint Conference on Neural Networks (IJCNN). New York: IEEE Press, 2013: 1-8. DOI:10.1109/IJCNN.2013.6706807 .
[31]
HEK M, ZHANGX Y, RENS Q, et al. Deep residual learning for image recognition[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition. New York: IEEE Press, 2016: 770-778. DOI:10.1109/CVPR.2016.90 .