In order to solve the problem of poor applicability of a single machine learning algorithm in data classification and prediction tasks in different fields, and to alleviate the impact of severe imbalance in datasets on prediction performance, a learning result prediction method based on Synthetic Minority Oversampling (SMOTE) and the ensemble model was proposed. The traditional SMOTE algorithm generated new synthetic samples by interpolating minority class samples, which could result in the presence of noise and high similarity between synthetic samples. To address these issues, an improved SMOTE algorithm was proposed, which removed noisy and easily confused samples by distance calculation, resulting in high discriminative and pure synthetic samples. Subsequently, an ensemble method was utilized to adjust the weights of samples and classifiers, leading to the creation of a stronger classifier with improved classification performance. Experimental results on the public online learning dataset Kalboard360 show that when using the Extreme Randomized Trees (ERT) classifier, in combination with improved SMOTE and Ensemble model, resulted in a prediction accuracy of 97.9%, which is a 5.5% increase compared to using a single ERT classifier. This demonstrates that the proposed SMOTE algorithm can generate high-quality balanced data, and the performance of the Ensemble learning model is significantly better than that of a single machine learning algorithm.
LIRuifeng, YANGHaifeng, CAIJianghui, et al.Outlier data mining algorithm based on weighted deep forest[J].Journal of Chinese Computer Systems, 2022, 43(7): 1426-1431.(in Chinese)
[3]
FISCHERC, PARDOSZ A, BAKERR S, et al.Mining big data in education: affordances and challenges[J].Review of Research in Education, 2020, 44(1): 130-160.
LIUTieyuan, CHENWei, CHANGLiang, et al.Research advances in the knowledge tracing based on deep learning[J].Journal of Computer Research and Development, 2022, 59(1): 81-104.(in Chinese)
CHENXi, MEIGuang, ZHANGJinjin, et al.Student grade prediction method based on knowledge graph and collaborative filtering[J].Journal of Computer Applications, 2020, 40(2): 595-601.(in Chinese)
[8]
TRAKUNPHUTTHIRAKR, CHEUNGY, LEEV C S.A study of educational data mining: Evidence from a thai university[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2019, 33(1): 734-741.
SHENHangjie, JUShenggen, SUNJieping. Performance prediction based on fuzzy clustering and support vector regression[J]. Journal of East China Normal University(Natural Science), 2019(5): 66-73.(in Chinese)
ZHANGYang, LUMingming, ZHENGYiji, et al.Student grade prediction based on graph auto-encoder model[J].Computer Engineering and Applications, 2021, 57(13): 251-257.(in Chinese)
[15]
DONGX, YUZ, CAOW, et al.A survey on ensemble learning[J].Frontiers of Computer Science, 2020, 14(1): 241-258.
LIHuifang, HUANGJianghang, XUGuanghao, et al.Multi-dimensional feature fusion-based runtime prediction approach for cloud workflow tasks[J].Acta Automatica Sinica, 2023, 49(1): 67-78.(in Chinese)
[18]
KAMALP, AHUJAS.An ensemble-based model for prediction of academic performance of students in undergrad professional course[J].Journal of Engineering, Design and Technology, 2019, 17(4): 769-781.
MEIDacheng, CHENJiang, ZHENGTao.Research on SMOTE algorithm based on boundary and density adaptation[J].Application Research of Computers, 2022, 39(5): 1478-1482.(in Chinese)
[21]
SALAWUM D, AROWOLOM O, ABDULSALAMS O, et al.A chi-square-SVM based pedagogical rule extraction method for microarray data analysis[J].International Journal of Advances in Applied Sciences, 2020, 9(2): 93-100.
[22]
CHENX, YUANY, ORGUNM A.Using Bayesian networks with hidden variables for identifying trustworthy users in social networks[J].Journal of Information Science, 2020, 46(5): 600-615.
[23]
YANGF J.An extended idea about decision trees[C]//2019 International Conference on Computational Science and Computational Intelligence (CSCI), 2019: 349-354.
[24]
ACOSTAM R C, AHMEDS, GARCIAC E, et al.Extremely randomized trees-based scheme for stealthy cyber-attack detection in smart grid networks[J].IEEE, 2020, 8: 19921-19933.
[25]
KURANIA, DOSHIP, VAKHARIAA, et al.A comprehensive comparative study of artificial neural network (ANN) and support vector machines (SVM) on stock forecasting[J].Annals of Data Science, 2023, 10(1): 183-208.
[26]
ASHRAFM, ZAMANM, AHMEDM.An intelligent prediction system for educational data mining based on ensemble and filtering approaches[J].Procedia Computer Science, 2020, 167(1): 1471-1483.
[27]
AMRIEHE A, HAMTINIT, ALJARAHI.Mining educational data to predict student’s academic performance using ensemble methods[J].International Journal of Database Theory and Application, 2016, 9(8): 119-136.
[28]
HODSONT O.Root-mean-square error (RMSE) or mean absolute error (MAE): when to use them or not[J].Geoscientific Model Development, 2022, 15(14): 5481-5487.