In order to identify anomalous behavior in encrypted network traffic while protecting data privacy, a privacy-preserving network traffic analysis scheme based on graph data augmentation (GDA) and contrastive learning is proposed.First, a provenance graph is constructed from network behavior logs to capture interactions and causal dependencies among nodes. Then, a differential-privacy-based graph data augmentation method is proposed, where the exponential mechanism is used to score and selectively perturb nodes and edges to generate privacy-preserving augmented views. Next, contrastive learning tasks are built on these augmented views to improve the consistency and discriminability of node representations. Finally, GraphSAGE is used to learn node representations and identify anomalous behaviors. Experimental results show that on the Streamspot dataset, the proposed scheme achieves both precision and accuracy as high as 99%, which are 25 percentage points and 33 percentage points higher than those of Streamspot method, respectively. The scheme can effectively identify malicious traffic in encrypted network traffic and shows greater advantages in privacy protection compared to baseline approaches.
针对上述挑战,本文引入图数据增强(Graph Data Augmentation,GDA)与对比学习技术。通过将网络行为抽象为数据溯源图结构并利用其内在关联进行分析,摆脱对数据包内容的依赖,从而缓解传统内容检测在加密流量下的性能退化问题;通过基于差分隐私的图数据增强方法对节点和边进行选择性扰动,生成多个隐私保护的增强视图,在扩展训练样本结构分布的同时降低敏感关系泄露风险;进一步利用对比学习约束不同增强视图下节点表示的一致性和判别性,从而缓解异常样本稀缺导致的模型偏倚问题。本文主要贡献如下:
ZHANGC L, FUY L, LIH, et al. Research on security scenarios and security models for 6G networking[J]. Chinese Journal of Network and Information Security, 2021, 7(1): 28-45. DOI:10.11959/j.issn.2096-109x.2021004(Ch ).
[3]
PAPADOGIANNAKIE, IOANNIDISS. A survey on encrypted network traffic analysis applications, techniques, and countermeasures[J]. ACM Computing Surveys, 2021, 54(6): 1-35. DOI:10.1145/3457904 .
FANGB X, JIAY, LIA P, et al. SARPPR: Reconstructing cyberspace security defense model[J]. Journal of Cybersecurity, 2024, 2(1): 2-12. DOI:10.20172/j.issn.2097-3136.240101(Ch ).
[8]
ROESCHM. Snort — lightweight intrusion detection for networks[C]// Proceedings of the 13th USENIX conference on System administration. New York: ACM. 1999: 229-238. DOI: 10.5555/1039834.1039864 .
[9]
PAXSONV. Bro: A system for detecting network intruders in real-time[J]. Computer Networks, 1999, 31(23-24): 2435-2463. DOI:10.1016/S1389-1286(99)00112-7 .
[10]
GARCÍAS, GRILLM, STIBOREKJ, et al. An empirical comparison of botnet detection methods[J]. Computers and Security, 2014, 45: 100-123. DOI:10.1016/j.cose.2014.05.011 .
[11]
LANSKYJ, ALI S, MOHAMMADIM, et al. Deep learning-based intrusion detection systems: A systematic review[J]. IEEE Access, 2021, 9: 101574-101599. DOI:10.1109/ACCESS.2021.3097247 .
[12]
LOTFOLLAHIM, JAFARI SIAVOSHANIM, SHIRALI HOSSEIN ZADER, et al. Deep packet: A novel approach for encrypted traffic classification using deep learning[J]. Soft Computing, 2020, 24(3): 1999-2012. DOI:10.1007/s00500-019-04030-2 .
[13]
SHAPIRAT, SHAVITTY. FlowPic: Encrypted Internet traffic classification is as easy as image recognition[C]//IEEE INFOCOM 2019 - IEEE Conference on Computer Communications Workshops (INFOCOM WKSHPS). New York: IEEE Press, 2019: 680-687. DOI:10.1109/INFCOMW.2019.8845315 .
[14]
WANGZ H, FOKK W, THINGV L L. Machine learning for encrypted malicious traffic detection: Approaches, datasets and comparative study[J]. Computers & Security, 2022, 113: 102542. DOI:10.1016/j.cose.2021.102542 .
[15]
HUOHT L, LUOY, LIP L, et al. Flow-based encrypted network traffic classification with graph neural networks[J]. IEEE Transactions on Network and Service Management, 2023, 20(2): 1224-1237. DOI:10.1109/TNSM.2022.3227500 .
[16]
MILAJERDIS M, GJOMEMOR, ESHETEB, et al. HOLMES: Real-time APT detection through correlation of suspicious information flows[C]//2019 IEEE Symposium on Security and Privacy (SP). New York: IEEE Press, 2019: 1137-1152. DOI:10.1109/SP.2019.00026 .
[17]
MILAJERDIS M, ESHETEB, GJOMEMOR, et al. POIROT: Aligning attack behavior with kernel audit records for cyber threat hunting[C]//Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security. New York: ACM, 2019: 1795-1812.. DOI:10.1145/3319535.3363217 .
[18]
MANZOORE, MILAJERDIS M, AKOGLUL. Fast memory-efficient anomaly detection in streaming heterogeneous graphs[C]//Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. New York: ACM, 2016: 1035-1044. DOI:10.1145/2939672.2939783 .
[19]
HANX Y, PASQUIERT, BATESA, et al. Unicorn: Runtime provenance-based detector for advanced persistent threats[C]//Proceedings 2020 Network and Distributed System Security Symposium. San Diego: Internet Society, 2020: 1-18. DOI:10.14722/ndss.2020.24046. DOI:10.14722/ndss.2020.24046 .
[20]
WANGQ, HASSANW U, LID, et al. You are what you do: Hunting stealthy malware via data provenance analysis[EB/OL]. [2024-09-13]. DOI: 10.14722/ndss.2020.24167 .
ZHUB J, LIN, CHENJ, et al. Encrypted traffic detection of domain generation algorithm based on contrastive learning[J]. Journal of Wuhan University (Natural Science Edition), 2025, 71(4): 517-525. DOI:10.14188/j.1671-8836.2024.0034(Ch ).
[23]
HAMILTONW L, YINGR, LESKOVECJ. Inductive representation learning on large graphs[EB/OL]. 2017: arXiv: 1706.02216. DOI: 10.48550/arXiv.1706.02216 .
[24]
DWORKC. Differential privacy[M]//Automata, Languages and Programming. Berlin: Springer, 2006: 1-12. DOI:10.1007/11787006_1 .
[25]
PASQUIERT, HANX Y, GOLDSTEINM, et al. Practical whole-system provenance capture[C]//Proceedings of the 2017 Symposium on Cloud Computing. New York: ACM, 2017: 405-418. DOI:10.1145/3127479.3129249 .
[26]
CHENT, KORNBLITHS, NOROUZIM, et al. A simple framework for contrastive learning of visual representations[EB/OL]. 2020: arXiv: 2002.05709.
[27]
HEK M, FANH Q, WUY X, et al. Momentum contrast for unsupervised visual representation learning[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). New York: IEEE Press, 2020: 9726-9735. DOI:10.1109/CVPR42600.2020.00975 .
[28]
GOYALA, WANGG, BATESA. R-CAID: Embedding root cause analysis within provenance-based intrusion detection[C]//2024 IEEE Symposium on Security and Privacy (SP). New York: IEEE Press, 2024: 3515-3532. DOI:10.1109/SP54263.2024.00253 .