1.Key Laboratory of Aerospace Information Security and Trusted Computing,Ministry of Education,School of Cyber Science and Engineering,Wuhan University,Wuhan 430072,Hubei,China
2.Wuhan Health Information Center,Wuhan 430014,Hubei,China
Show less
文章历史+
Received
Published
2024-03-04
2025-08-24
Issue Date
2026-07-23
PDF (2598K)
摘要
为进一步提升在加密DNS(Domain Name System)的通讯场景中,DGA(Domain Generation Algorithm)流量检测的实时性、准确率和泛化性,保障计算机通信安全,提出了一种域名生成算法加密流量检测方案——EDGAD(Encrypted Domain Generation Algorithm Detection)。EDGAD由数据预处理和流量分类两个模块构成。预处理模块通过流量簇划分策略和高效特征集减小开销。流量分类模块包含两阶段分类模型,其中阶段一采用二分类模型将DoH(DNS over HTTPS)流量分为DGA流量和非DGA流量;阶段二基于对比学习思想构建多分类模型,识别DGA加密流量所属的具体DGA软件。此外,阶段二中还设计了流量样本增强方法,以及对比学习模块和多分类模块联合调优方法,以提升模型泛化性和训练效率。在采用1 s流量数据构建特征的数据集实验中,选择XGBoost模型作为阶段一的二分类模型,并将阶段二中缩放参数确定为0.25。实验结果表明,EDGAD能有效识别7种DGA恶意软件,准确率达到98.07%,平均精度均值达到0.981 3,较对比方案分别提升了1.24个百分点和0.013 6。
Abstract
An encrypted traffic detection scheme (EDGAD) is proposed to enhance the real-time detection, accuracy, and generalization of domain generation algorithm (DGA) traffic in encrypted domain name system (DNS) communications and ensure computer communication security. EDGAD comprises two modules: data preprocessing and traffic classification. The preprocessing module employs traffic clustering and an efficient feature set to reduce the overhead. The traffic classification module includes a two-stage model. In the first stage, a binary classifier is used to distinguish DNS from HTTPS (DoH) traffic as either DGA or non-DGA. The second stage employs contrastive learning to construct a multiclass model that identifies the specific DGA software that generates the encrypted traffic. This stage also incorporates traffic sample enhancement and the joint optimization of contrastive learning and multi-classification modules to improve generalization and training efficiency. Experiments using 1 s traffic data indicated that the XGBoost model was optimal for the first-stage binary classification, with the second-stage scaling parameter set to 0.25. The results demonstrate that EDGAD effectively identifies seven types of DGA malware, with an accuracy of 98.07%, and an average precision mean value of 0.981 3, and an improvement of 1.24 percentage points and 0.013 6 compared with the comparison scheme, respectively.
传统域名系统(Domain Name System, DNS)以明文形式进行域名查询[1],带来了隐私泄露风险[2]。为此,服务商开始支持加密DNS协议,如DNS over HTTPS(DoH)[3]和DNS over TLS(DoT)[4],以保护DNS消息的完整性和用户隐私。然而,这些加密协议也可能被攻击者用来隐藏恶意软件的通信内容,使得传统基于明文流量的检测难以识别其恶意活动,从而逃避网络监管[5]。例如域名生成算法(Domain Generation Algorithm, DGA)恶意软件通过生成大量随机域名规避传统安全检测,从而使受感染主机能够连接到僵尸网络。为了应对这一威胁,研究人员开发了多种技术,包括通过逆向工程[6]识别DGA背后的家族和变种,以及通过监控内网DNS流量[7]并维护白名单数据库来识别和阻断异常DNS流量。
SCHMIDG. Thirty years of DNS insecurity: Current issues and perspectives[J]. IEEE Communications Surveys & Tutorials, 2021, 23(4): 2429-2459. DOI: 10.1109/COMST.2021.3105741 .
[2]
GARCÍAS, HYNEKK, VEKSHIND, et al. Large scale measurement on the adoption of encrypted DNS[EB/OL]. 2021: arXiv: 2107.04436.
[3]
HOFFMANP, MCMANUSP. DNS queries over HTTPS (DoH)[J]. RFC, 2018, 8484: 1-21. DOI: 10.17487/RFC8484 .
[4]
LYUM Z, GHARAKHEILIH H, SIVARAMANV. A survey on DNS encryption: Current development, malware misuse, and inference techniques[J]. ACM Computing Surveys, 2023, 55(8): 1-28. DOI: 10.1145/3547331 .
[5]
PATSAKISC, CASINOF, KATOSV. Encrypted and covert DNS queries for botnets: Challenges and countermeasures[J]. Computers & Security, 2020, 88: 101614. DOI: 10.1016/j.cose.2019.101614 .
[6]
PLOHMANND, YAKDANK, KLATTM, et al. A comprehensive measurement study of domain generating malware[C]//Proceedings of the 25th USENIX Conference on Security Symposium. Austin: USENIX Association, 2016: 263–278.
[7]
ICHISEH, JINY, IIDAK, et al. NS record history based abnormal DNS traffic detection considering adaptive botnet communication blocking[J]. Journal of Information Processing, 2020, 28: 112-122. DOI: 10.2197/ipsjjip.28.112 .
[8]
MITSUHASHIR, JINY, IIDAK, et al. Detection of DGA-based malware communications from DoH traffic using machine learning analysis[C]//2023 IEEE 20th Consumer Communications & Networking Conference (CCNC). New York: IEEE Press, 2023: 224-229. DOI: 10.1109/CCNC51644.2023.10059835 .
[9]
CHENT, KORNBLITHS, NOROUZIM, et al. A simple framework for contrastive learning of visual representations[EB/OL]. [2020-03-30].
FANGH C, WANGS C, ZHOUM, et al. CERT: Contrastive self-supervised learning for language understanding[EB/OL]. 2020: arXiv: 2005.12766.
[12]
KHOSLAP, TETERWAKP, WANGC, et al. Supervised contrastive learning[EB/OL]. [2020-04-23].
[13]
ZHAOH, CHANGZ B, BAOG B, et al. Malicious domain names detection algorithm based on N-gram[J]. Journal of Computer Networks and Communications, 2019, 2019(1): 4612474.1-4612474.9. DOI: 10.1155/2019/4612474 .
[14]
SURYOTRISONGKOH, MUSASHIY, TSUNEDAA, et al. Robust botnet DGA detection: Blending XAI and OSINT for cyber threat intelligence sharing[J]. IEEE Access, 2022, 10: 34613-34624. DOI: 10.1109/ACCESS.2022.3162588 .
[15]
SCHÜPPENS, TEUBERTD, HERRMANNP, et al. FANCI: Feature-based automated NXDomain classification and intelligence[C]//Proceedings of the 27th USENIX Conference on Security Symposium. Baltimore: USENIX Association, 2018: 1165-1181. DOI: 10.5555/3277203.3277290 .
[16]
PATSAKISC, CASINOF. Exploiting statistical and structural features for the detection of Domain Generation Algorithms[J]. Journal of Information Security and Applications, 2021, 58: 102725. DOI: 10.1016/j.jisa.2020.102725 .
[17]
ZHANGY D, CHENY Z, LINY Y, et al. Detection of algorithmically generated domain names using SMOTE and hybrid neural network[C]//CCF Conference on Computer Supported Cooperative Work and Social Computing. Singapore: Springer, 2019: 738-751.10.1007/978-981-15-1377-0_57. DOI: 10.1007/978-981-15-1377-0_57 .
[18]
AYUBM A, SMITHS, SIRAJA, et al. Domain generating algorithm based malicious domains detection[C]//2021 8th IEEE International Conference on Cyber Security and Cloud Computing (CSCloud)/2021 7th IEEE International Conference on Edge Computing and Scalable Cloud (EdgeCom). New York: IEEE Press, 2021: 77-82. DOI: 10.1109/CSCloud-EdgeCom52276.2021.00024 .
[19]
RAVIV, ALAZABM, SRINIVASANS, et al. Adversarial defense: DGA-based botnets and DNS homographs detection through integrated deep learning[J]. IEEE Transactions on Engineering Management, 2023, 70(1): 249-266. DOI: 10.1109/TEM.2021.3059664 .
[20]
VEKSHIND, HYNEKK, CEJKAT. DoH Insight: Detecting DNS over HTTPS by machine learning[C]//Proceedings of the 15th International Conference on Availability, Reliability and Security. New York: ACM, 2020: 87.1-87.8. DOI: 10.1145/3407023.3409192 .
[21]
CSIKORL, SINGHH, KANGM S, et al. Privacy of DNS-over-HTTPS: Requiem for a dream?[C]//2021 IEEE European Symposium on Security and Privacy (EuroS&P). New York: IEEE Press, 2021: 252-271. DOI: 10.1109/EuroSP51992.2021.00026 .
[22]
SIBYS, JUAREZM, DIAZC, et al. Encrypted DNS ⇒ Privacy? A traffic analysis perspective[EB/OL]. 2019: arXiv: 1906.09682. DOI: 10.14722/ndss.2020.24301 .
[23]
BUSHARTJ, ROSSOWC. Padding ain’t enough: Assessing the privacy guarantees of encrypted DNS[EB/OL]. 2019: arXiv: 1907.01317.
Canadian Institute for Cybersecurity. CIRA-CIC-DoHBrw-2020[DB/OL]. [2024-03-01].
[26]
MONTAZERISHATOORIM, DAVIDSONL, KAURG, et al. Detection of DoH tunnels using time-series classification of encrypted traffic[C]//2020 IEEE Intl Conf on Dependable, Autonomic and Secure Computing, Intl Conf on Pervasive Intelligence and Computing, Intl Conf on Cloud and Big Data Computing, Intl Conf on Cyber Science and Technology Congress(DASC/PiCom/CBDCom/CyberSciTech). New York: IEEE Press, 2020: 63-70. DOI: 10.1109/DASC-PICom-CBDCom-CyberSciTech49142.2020.00026 .