To address the issues of temporal feature loss caused by traffic visualization and the insufficient detection accuracy of lightweight models under complex attacks in network intrusion detection, a lightweight model named MobileViT-AESTF was proposed based on recurrence plot (RP) and an adaptive efficient spatiotemporal fusion (AESTF) module. Utilizing phase space reconstruction technology, the method mapped one-dimensional traffic time series into recursive relationships in a two-dimensional state space via RP, effectively retaining the non-linear temporal features of the traffic. To overcome the insufficient feature capture capability of MobileViT_V3 in processing visualized traffic, the designed AESTF module comprised an efficient spatiotemporal fusion (ESTF) unit and an adaptive feature exciter (AFE); specifically, the ESTF employed a dual-path heterogeneous attention mechanism to extract spatial and channel features in parallel, while the AFE adaptively enhanced key discriminative features based on a neuronal energy function. Experimental results on the CIC UNSW-NB15 Augmented dataset demonstrate that MobileViT-AESTF achieves an accuracy, recall, and F1-score of 99.84%, 99.82%, and 99.85% respectively, while maintaining a low parameter count of 1.04×106. Compared with the baseline MobileViT_V3, the proposed model reduces computational cost and exhibits greater robustness in identifying complex attacks such as Fuzzers. The model achieves an effective balance between lightweight design and detection accuracy, making it suitable for network intrusion detection in resource-constrained environments and offering significant engineering application value.
BlackD. Cybercrime will cost $12TN next year, say experts [EB/OL]. (2024-01-24)[2025-08-04].
[2]
LiaoH, MurahM Z, HasanM K, et al. A survey of deep learning technologies for intrusion detection in internet of things[J]. IEEE Access, 2024, 12: 4745-4761.
[3]
HanY L, HanX W. ICPS multi-target constrained comprehensive security control based on DoS attacks energy grading detection and compensation[J]. Journal of Measurement Science and Instrumentation, 2024, 15(4): 518-531.
[4]
DemmeseF A, NeupaneA, KhorsandrooS, et al.Machine learning based fileless malware traffic classification using image visualization[J].Cybersecurity, 2023, 6(1): 1-18.
LiuWenqi, HuTao, YanJie, et al. Network intrusion detection technology based on DeepInsight and transfer learning[J]. Chinese Journal of Engineering, 2024, 46(12): 2238-2245.(in Chinese)
[7]
LiS M, ChaiG Z, WangY H, et al. CRSF: an intrusion detection framework for industrial internet of things based on pretrained CNN2D-RNN and SVM[J]. IEEE Access, 2023, 11: 92041-92054.
[8]
PhamV, SeoE, ChungT M. Lightweight convolutional neural network based intrusion detection system[J]. Journal of Communications, 2020, 15(11): 808-817.
YaoJun, SunFangchao. Research on lightweight intrusion detection model based on MobileViT[J]. Modern Electronics Technique, 2024, 47(19): 33-39.(in Chinese)
[11]
TiwariR S, LakshmiD, DasT K, et al. A lightweight optimized intrusion detection system using machine learning for edge-based IIoT security[J]. Telecommunication Systems, 2024, 87(3): 605-624.
[12]
AlshehriM S, SaidaniO, AlrayesF S, et al. A self-attention-based deep convolutional neural networks for IIoT networks intrusion detection[J]. IEEE Access, 2024, 12: 45762-45772.
[13]
Al-JarrahO Y, El HalouiK, DianatiM, et al. A novel detection approach of unknown cyber-attacks for intra-vehicle networks using recurrence plots and neural networks[J]. IEEE Open Journal of Vehicular Technology, 2023, 4: 271-280.
[14]
WadekarS N, ChaurasiaA, MishraA. MobileViTv3: enhanced fusion for mobile vision transformers[C]//Proceedings of the European Conference on Computer Vision, 2023: 412-429.
[15]
YangL X, ZhangR Y, LiL D, et al. SimAM: A simple, parameter-free attention module for convolutional neural networks[C]//38th International Conference on Machine Learning, 2021: 11863-11874.
[16]
WangY, LiY S, WangG, et al. Multi-scale attention network for single image super-resolution[C]//IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2024: 5950-5960.
[17]
Canadian Institute for Cybersecurity (CIC), University of New Brunswick. CIC UNSW-NB15 Augmented Dataset[DB/OL]. [2025-09-07].
[18]
MaN N, ZhangX Y, ZhengH T, et al. Shufflenet v2: Practical guidelines for efficient cnn architecture design[C]//European Conference on Computer Vision (ECCV), 2018: 116-131.
[19]
ChenX, LiuZ, WangY, et al. Efficient mobile transformer for edge-AI security applications[C]//International Conference on Machine Learning, 2023: 2689-2699.
[20]
HowardA G, ZhuM, ChenB, et al. Mobilenets: efficient convolutional neural networks for mobile vision applications[PP/OL].Vl.arXiv(2017-04-17)[2025-09-07].
[21]
SandlerM, HowardA, ZhuM L, et al. MobileNetV2: Inverted residuals and linear bottlenecks[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018: 4510-4520.
[22]
HowardA, SandlerM, ChenB, et al. Searching for MobileNetV3[C]//2019 IEEE/CVF International Conference on Computer Vision(ICCV), 2019: 1314-1324.
[23]
MehtaS, RastegariM. Mobilevit: light-weight, general-purpose, and mobile-friendly vision transformer[PP/OL].Vl.arXiv(2021-10-05)[2025-09-07].
[24]
MehtaS, RastegariM. Separable self-attention for mobile vision transformers[PP/OL].Vl.arXiv(2022-06-06)[2025-09-07].
[25]
VasuP K A, GabrielJ, ZhuJ, et al. FastViT: a fast hybrid vision transformer using structural reparameterization[C]//IEEE/CVF International Conference on Computer Vision(ICCV), 2024: 5785-5794.
[26]
LuX Y, SuganumaM, OkataniT. SBCFormer: lightweight network capable of full-size imagenet classification at 1 fps on single board computers[C]//IEEE/CVF Winter Conference on Applications of Computer Vision, 2024: 1123-1133.
[27]
LiuC, ZhangH, WangZ, et al. Squeeze-attention: Lightweight network for real-time intrusion detection[J]. IEEE Transactions on Dependable and Secure Computing, 2024, 21(3): 1987-1999.