数据增强下深度卷积网络不变性机理的可解释性研究

刘妮 ,  雷娇娇 ,  邢振榕 ,  韩文静 ,  何立火

西安交通大学学报 ›› 2026, Vol. 60 ›› Issue (9) : 218 -228.

PDF (15823KB)
西安交通大学学报 ›› 2026, Vol. 60 ›› Issue (9) : 218 -228. DOI: 10.7652/xjtuxb202609021

数据增强下深度卷积网络不变性机理的可解释性研究

作者信息 +

Interpretability Study of the Invariance Mechanism of Deep Convolutional Networks under Data Augmentation

Author information +
文章历史 +
PDF (16202K)

摘要

针对目前数据增强提升深度卷积网络不变性的研究缺乏对其实现不变性的机理进行阐释的问题,对其位置、方向等信息的表征机制和特征综合机理展开研究。首先,采用圆形和方形二分类任务构建简化实验场景,使用其增强数据训练网络;然后,通过最大激活方法对全连接层和全局平均池化层进行特征可视化,并采用t-SNE降维揭示了全连接层间高维权重向量的分类决策机理。结果表明:全连接层首层通过逐步扩大位置感知域的方式实现平移不变性;在圆方数据集平移增强实验中,Fc2层平移不变性得分为0.99,较Conv5层提升了45%;全连接层通过线性加权机制综合圆形和方形的局部特征,遵循同类别特性正向激发、异类别特征负向抑制的连接逻辑;ResNet架构中平均池化层展现更优的平移不变性,在圆方数据集增强实验中,平均池化层得分为0.93,高于全连接层的0.79,而级联全连接结构则表现出更强的旋转和尺度不变性。以上结论在MNIST数据集的扩展实验上也得到了验证。

Abstract

Addressing the issue that current research on enhancing the invariance of deep convolutional networks through data augmentation lacks an explanation of the mechanisms underlying its implementation,this paper explores the representation and pooling mechanisms of position,orientation,and other information.First,a simplified experimental scenario is constructed using a binary classification task of circles and squares,which is employed to augment data for network training.Then,feature visualization is conducted for the fully connected layer and global average pooling layer using the maximum activation method,and t-SNE dimensionality reduction is adopted to reveal the classification decision-making mechanism of high-dimensional weight vectors within the fully connected layers.It is shown that translation invariance is achieved by the first fully connected layer through the gradual expansion of the position-aware receptive field.In the translation augmentation experiment on the circle-and-square dataset,a translation invariance score of 0.99 is obtained by the Fc2 layer,representing a 45% improvement over that of the Conv5 layer.The local features of circles and squares are synthesized by the fully connected layers through a linear weighting mechanism,following a connection logic in which same-category features are positively activated and different-category features are negatively suppressed.Better translation invariance is exhibited by the average pooling layer in the ResNet architecture.In the augmentation experiment on the circle-and-square dataset,a score of 0.93 is obtained by the AvgPool layer,higher than the 0.79 obtained by the fully connected layer,whereas stronger rotation and scale invariance is exhibited by the cascaded fully connected structure.These conclusions are also verified through extended experiments on the MNIST dataset.

关键词

深度卷积网络 / 数据增强 / 平移不变性 / 特征可视化

Key words

deep convolutional network / data augmentation / translation invariance / feature visualization

引用本文

引用格式 ▾
刘妮,雷娇娇,邢振榕,韩文静,何立火. 数据增强下深度卷积网络不变性机理的可解释性研究[J]. 西安交通大学学报, 2026, 60(9): 218-228 DOI:10.7652/xjtuxb202609021

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1]

Blything R,Biscione V,Vankov I I,et al. The human visual system and CNNs can both support robust online translation tolerance following extreme displacements[J]. Journal of Vision, 2021,21(2):9.

[2]

Jaderberg M,Simonyan K,Zisserman A,et al. Spatial transformer networks[C]//Proceedings of the 28th International Conference on Neural Information Processing Systems.Cambridge,MA,USA:MIT Press,2015:2017-2025.

[3]

Biscione V,Bowers J.Learning translation invariance in CNNs[PP/OL].arXiv(2020—11—06)[2025—10—13].https://arxiv.org/abs/2011.11757.

[4]

Laptev D,Savinov N,Buhmann J M,et al. TI—POOLING:transformation—invariant pooling for feature learning in convolutional neural networks[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR).Piscataway,NJ,USA:IEEE,2016:289-297.

[5]

顾正强,张严,张冰尘. 一种基于模糊滤波提高SAR自动目标识别平移不变性的方法[J]. 系统工程与电子技术, 2020,42(11):2488-2496.

[6]

Gu Zhengqiang,Zhang Yan,Zhang Bingchen. Improving translation invariance of SAR automatic target recognition based on blur filtering method[J]. Systems Engineering and Electronics, 2020,42(11):2488-2496.

[7]

李俊英. 深度卷积神经网络的旋转等变性研究[D].杭州:浙江大学,2019.

[8]

Krizhevsky A,Sutskever I,Hinton G E. ImageNet classification with deep convolutional neural networks[C]//Proceedings of the 26th International Conference on Neural Information Processing Systems.Red Hook,NY,USA:Curran Associates Inc.,2012:1097-1105.

[9]

Hermann K L,Chen Ting,Kornblith S. The origins and prevalence of texture bias in convolutional neural networks[C]//Proceedings of the 34th International Conference on Neural Information Processing Systems.Red Hook,NY,USA:Curran Associates Inc.,2020:19000-19015.

[10]

Myburgh J C,Mouton C,Davel M H. Tracking translation invariance in CNNs[C]//Artificial Intelligence Research.Cham,Switzerland:Springer International Publishing,2020:282-295.

[11]

Kauderer—Abrams E. Quantifying translation—invariance in convolutional neural networks[PP/OL].arXiv(2017—12—10)[2025—10—13].https://arxiv.org/abs/1801.01450.

[12]

Baker N,Lu Hongjing,Erlikhman G,et al. Local features and global shape information in object classification by deep convolutional neural networks[J]. Vision Research, 2020,172:46-61.

[13]

Erhan D,Bengio Y,Courville A,et al. Visualizing higher—layer features of a deep network[J]. University of Montreal, 2009,1341(3):1.

[14]

Yosinski J,Clune J,Nguyen A,et al. Understanding neural networks through deep visualization[PP/OL].arXiv(2015—06—22)[2025—10—13].https://arxiv.org/abs/1506.06579.

[15]

Nguyen A,Yosinski J,Clune J.Understanding neural networks via feature visualization:a survey[M]//Samek W,Montavon G,Vedaldi A,et al. Explainable AI:Interpreting,Explaining and Visualizing Deep Learning.Cham,Switzerland:Springer International Publishing,2019:55-76.

[16]

Ribeiro M T,Singh S,Guestrin C. “Why should I trust you?”:explaining the predictions of any classifier[C]//Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining.New York,USA:ACM,2016:1135-1144.

[17]

van der Maaten L,Hinton G. Visualizing data using t—SNE[J]. Journal of Machine Learning Research, 2008,9(86):2579-2605.

[18]

Deng Li. The MNIST database of handwritten digit images for machine learning research[J]. IEEE Signal Processing Magazine, 2012,29(6):141-142.

[19]

Semih K O,van Gemert J C . On translation invariance in CNNs:convolutional layers can exploit absolute spatial location[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).Piscataway,NJ,USA:IEEE,2020:14262-14273.

[20]

Han Yena,Roig G,Geiger G,et al. Scale and translation—invariance for novel objects in human vision[J]. Scientific Reports, 2020,10(1):1411.

[21]

Zeiler M D,Fergus R. Visualizing and understanding convolutional networks[C]//Computer Vision—ECCV 2014.Cham,Switzerland:Springer International Publishing,2014:818-833.

[22]

Zhou Bolei,Khosla A,Lapedriza A,et al. Learning deep features for discriminative localization[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR).Piscataway,NJ,USA:IEEE,2016:2921-2929.

[23]

Zhang Quanshi,Yang Yu,Ma Haotian,et al. Interpreting CNNs via decision trees[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).Piscataway,NJ,USA:IEEE,2019:6254-6263.

[24]

Hu Jiacong,Gao Jing,Ye Jingwen,et al. Model LEGO:creating models like disassembling and assembling building blocks[C]//Proceedings of the 38th International Conference on Neural Information Processing Systems.Red Hook,NY,USA:Curran Associates Inc.,2024:127711-127738.

[25]

Nguyen A,Yosinski J,Clune J. Deep neural networks are easily fooled:high confidence predictions for unrecognizable images[C]//2015 IEEE Conference on Computer Vision and Pattern Recognition(CVPR).Piscataway,NJ,USA:IEEE,2015:427-436.

[26]

Szegedy C,Zaremba W,Sutskever I,et al. Intriguing properties of neural networks[PP/OL].arXiv(2014—02—19)[2025—10—13].https://arxiv.org/abs/1312.6199.

[27]

Wang Feng,Liu Haijun,Cheng Jian. Visualizing deep neural network by alternately image blurring and deblurring[J]. Neural Networks, 2018,97:162-172.

[28]

王芳珍,张小丽,赵琦武,. 一维卷积神经网络在机械故障特征提取中的可解释性研究[J]. 西安交通大学学报, 2025,59(7):24-35.

[29]

Wang Fangzhen,Zhang Xiaoli,Zhao Qiwu,et al. Study on the interpretability of one—dimensional convolutional neural networks in mechanical fault feature extraction[J]. Journal of Xi’an Jiaotong University, 2025,59(7):24-35.

[30]

Rousseeuw P J. Silhouettes:a graphical aid to the interpretation and validation of cluster analysis[J]. Journal of Computational and Applied Mathematics, 1987,20:53-65.

[31]

Jansson Y,Lindeberg T. Scale—invariant scale—channel networks:deep networks that generalise to previously unseen scales[J]. Journal of Mathematical Imaging and Vision, 2022,64(5):506-536.

基金资助

国家自然科学基金区域创新发展联合基金资助项目(U23A20682)

国家自然科学基金区域创新发展联合基金资助项目(U22A20247)

陕西省自然科学基础研究计划资助项目(2024JC-YBMS-566)

AI Summary AI Mindmap
PDF (15823KB)

0

访问

0

被引

详细

导航
相关文章

AI思维导图

/