A lightweight network intrusion detection model, termed MobileViT-RMF, was proposed to address several limitations in existing traffic-visualization-based intrusion detection methods, including insufficient representation of temporal correlations in traffic data, limited feature characterization capability, and restricted detection accuracy under lightweight model constraints. The proposed method combined the Gramian Angular Difference Field (GADF) with a Residual Multi-grain Focus (RMF) module to improve both traffic representation and feature extraction. Firstly, one-dimensional network traffic feature sequences were converted into two-dimensional GADF images with spatiotemporal correlations, so that the evolutionary characteristics of traffic behavior and the discriminative differences among traffic categories could be more effectively preserved in the image space. On this basis, the RMF module was embedded into the lightweight MobileViT backbone to strengthen local key-feature extraction and global contextual modeling through a multi-granularity feature focusing mechanism, thereby the discriminative capability of the model for complex traffic images was enhanced. Experimental validation was conducted on the CICIDS2017 and UNSW-NB15 public datasets. The proposed model achieves a detection accuracy of 99.75% on CICIDS2017, which is 2.74 percentage points higher than that of the baseline MobileViT_V3_xxs model. On UNSW-NB15, the peak accuracy reaches 99.89%, indicating favorable cross-dataset detection performance. These findings suggest that the combination of GADF-based traffic image representation and RMF-based feature enhancement can improve the representation quality and classification performance of network traffic images while preserving the lightweight nature of the model. The proposed approach provides a feasible idea for the design of lightweight network intrusion detection models.
本研究采用加拿大网络安全研究所发布的CICIDS2017基准数据集, 该数据集通过模拟真实企业网络环境而构建, 包含连续5 d的完整网络流量记录, 涵盖HTTP、 FTP、 SSH等常规协议流量及多种新型攻击模式[17]。实验选取BENIGN正常流量与DDoS、 DoS GoldenEye、 DoS Hulk、 DoS Slowhttptest、 DoS Slowloris五类典型DoS攻击数据, 该分类体系可有效验证模型在泛洪攻击DDoS、 应用层攻击Slowloris及混合攻击场景下的检测鲁棒性。
KUMARIP, JAINA K. A comprehensive study of DDoS attacks over IoT network and their countermeasures[J]. Computers & Security, 2023, 127: 103096.
[2]
NEXUSGUARD. Distributed denial of service trend report 2024[EB/OL].(2025-02-21)[2025-06-12].
[3]
ALI A, YOUSAFM M. Novel three-tier intrusion detection and prevention system in software defined network[J]. IEEE Access, 2020, 8: 109662-109676.
[4]
LIAOH, MURAHM Z, HASANM K, et al. A survey of deep learning technologies for intrusion detection in Internet of Things[J]. IEEE Access, 2024, 12: 4745-4761.
[5]
MOY M, LIH, WANGD S, et al. An intrusion detection system based on convolution neural network[J]. PeerJ Computer Science, 2024, 10: e2152.
[6]
DEMMESEF A, NEUPANEA, KHORSANDROOS, et al. Machine learning based fileless malware traffic classification using image visualization[J]. Cybersecurity, 2023, 6(1): 32.
[7]
PHAMV, SEOE, CHUNGT M. Lightweight convolutional neural network based intrusion detection system[J]. Journal of Communications, 2020, 15(11): 808-817.
[8]
ATTACKW, ATTACKI, ATTACKBF. Ensemble of feature augmented convolutional neural network and deep autoencoder for efficient detection of network attacks[J]. Scientific Reports, 2025, 15: 4267.
LIUWenqi, HUTao, YANJie, et al. Network intrusion detection technology based on DeepInsight and transfer learning[J]. Chinese Journal of Engineering, 2024, 46(12): 2238-2245.(in Chinese)
[11]
WUZ T, LONGZ H. Research on network traffic classification method based on CNN-RNN[M]. Singapore: Springer Nature, 2023.
[12]
WANGL T, HUW, LIUJ Y, et al. Encrypted traffic classification based on fusion of vision transformer and temporal features[J]. The Journal of China Universities of Posts and Telecommunications, 2023, 30(2): 73-82.
[13]
AVCIC, TEKINERDOGANB, CATALC. Design tactics for tailoring transformer architectures to cybersecurity challenges[J]. Cluster Computing, 2024, 27(7): 9587-9613.
YAOJun, SUNFangchao.Research on lightweight intrusion detection model based on MobileViT[J].Modern Electronics Technique, 2024,47(19): 33-39.(in Chinese)
[16]
XIAOF Y, CHENY Y, ZHUY H. GADF/GASF-HOG: Feature extraction methods for hand movement classification from surface electromyography[J]. Journal of Neural Engineering, 2020, 17(4): 046016.
[17]
WANGY, LIY S, WANGG, et al. Multi-scale attention network for single image super-resolution[C]//2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops(CVPRW), 2024: 5950-5960.
[18]
LUOS S, LIM J, CHENM, Multi-scale integrated at-tention mechanism for facial expression recognition network[J].Computer Engineering and Applications,2023,59(1): 199-206.
[19]
PANIGRANHIR, BORAHS. A detailed analysis of CICIDS2017 dataset for designing intrusion detection systems[J]. International Journal of Engineering & Technology, 2018, 7(3): 479-482.
[20]
PRADIPTAG A, WARDOYOR, MUSDHOLIFAHA, et al. SMOTE for handling imbalanced data problem: A review[C]//2021 Sixth International Conference on Informatics and Computing (ICIC), 2021: 1-8.
[21]
PEDREGOSAF, VAROQUAUXG, GRAMFORTA, et al. Scikit-learn: Machine learning in Python[J]. Journal of Machine Learning Research, 2011, 12(10): 2825-2830.
[22]
ZHUANGZ X, LIUM R, CUTKOSKYA, et al. Understanding AdamW through proximal methods and scale-freeness[DB/OL]. (2022-01-31)[2025-06-16].
[23]
MAN N, ZHANGX Y, ZHENGH T, et al. Shufflenet V2: Practical guidelines for efficient CNN architecture design[DB/OL]. (2018-07-30)[2025-06-12].
[24]
BODAVARAPUP N R, SRINIVASP V V S. Facial expression recognition for low resolution images using convolutional neural networks and denoising techniques[J]. Indian Journal of Science and Technology,2021,14(12): 971-983.
[25]
HOWARDA G, ZHUM, CHENB, et al. Mobilenets: Efficient convolutional neural networks for mobile vision applications[DB/OL]. (2017-04-17)[2025-06-12].
[26]
SANDLERM, HOWARDA, ZHUM, et al. MobileNetV2: Inverted residuals and linear bottlenecks[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018: 4510-4520.
[27]
HOWARDA, SANDLERM, CHUB, et al. Searching for MobileNetV3[C]//2019 IEEE/CVF International Conference on Computer Vision(ICCV), 2019: 1314-1324.
[28]
MEHTAS, RASTEGARIM. Mobilevit: Light-weight, general-purpose, and mobile-friendly vision transformer[DB/OL]. (2021-10-05)[2025-06-12].
[29]
MEHTAS, RASTEGARIM. Separable self-attention for mobile vision transformers[DB/OL]. (2022-06-06)[2025-06-12].
[30]
WADEKARS N, CHAURASIAA. Mobilevitv3: Mobile-friendly vision transformer with simple and effective fusion of local, global and input features[DB/OL]. (2022-09-30)[2025-06-12].
[31]
IANDOLAF N, HANS, MOSKEWICZM W, et al. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5 MB model size[DB/OL].(2016-02-24)[2025-06-12]. . org/abs/1602. 07360.
[32]
VERKERKENM, D’HOOGEL, SUDYANAD, et al. A novel multi-stage approach for hierarchical intrusion detection[J]. IEEE Transactions on Network and Service Management, 2023, 20(3): 3915-3929.
[33]
GORYUNOVM N, MATSKEVICHA G, RYBOLOVLEVD A. Synthesis of a machine learning model for detecting computer attacks based on the CICIDS2017 dataset[J]. Proceedings of the Institute for System Programming of the RAS, 2020, 32(5): 81-94.
[34]
GETMANA I, RYBOLOVLEVD A, NIKOLSKAYAA G. Deep learning applications for intrusion detection in network traffic[J]. Programming and Computer Software, 2024, 50(7): 493-510.
[35]
BELARBIO, KHANA, CARNELLIP, et al. An intrusion detection system based on deep belief networks[C]//International Conference on Science of Cyber Security. Cham: Springer International Publishing, 2022: 377-392.
[36]
MOUSTAFAN, SLAYJ. UNSW-NB15: A comprehensive data set for network intrusion detection systems (UNSW-NB15 network data set)[C]//2015 Military Communications and Information Systems Conference (MilCIS), 2015: 1-6.