To solve the problem that the classic K-means algorithm divides blob dataset into cluster structure with uncertainty and may cause distortion of clustering result, this paper proposes a set of methods for determining the number of clusters K, the cluster centers and cluster boundaries. Based on the sample mean and standard deviation of the data features, the maximum number of clusters Kmax of the dataset is estimated using the prior rule. For different numbers within range of K∈[1,Kmax], the multiple clustering performance indicators are calculated, and the inherent number of clusters Kbest is determined by finding the maximum value of each indicator. With Kbest as the input parameter of the K-means algorithm, the corresponding cluster centers are iteratively updated using the expectation-maximization (E-M) algorithm. The circular boundaries for each cluster are defined with the Euclidean distance from the farthest data point in the same cluster to the cluster center as the radius. The validity and robustness of the proposed methods are verified based on the simulation experiment of two-dimensional blob random data with different sample sizes and different number of clusters.
BANDARUS, NGA H C, DEBK.Data mining methods for knowledge discovery in multi-objective optimization:part A-Survey[J].Expert Systems with Applications,2017,70:139-159.
[2]
BUIA T.Dimension reduction with prior information for knowledge discovery[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2024,46(5):3625-3636.
[3]
CZARNOWSKII, JĘDRZEJOWICZP.Supervised classification problems-taxonomy of dimensions and notation for problems identification[J].IEEE Access,2021,9:151386-151400.
[4]
SINGHJ, SINGHD.A comprehensive review of clustering techniques in artificial intelligence for knowledge discovery:taxonomy,challenges, applications and future prospects[J].Advanced Engineering Informatics, 2024,62:102799.
[5]
MAHNOOR, SHAFII, CHAUDHRYM,et al.A review of approaches for rapid data clustering:challenges, opportunities,and future directions [J].IEEE Access,2024,12:138086-138120.
[6]
SUBASIA.Machine learning techniques[M]//Practical Machine Learning for Data Analysis Using Python.Amsterdam:Elsevier,2020: 91-202.
STEINLEYD. K-means clustering:a half-century synthesis[J].British Journal of Mathematical and Statistical Psychology,2006,59:1-34.
[9]
STEINLEYD.Stability analysis in K-means clustering[J].British Journal of Mathematical and Statistical Psychology,2008,61:255-273.
[10]
SELIMS Z, ISMAILM A. K-means-type algorithms:a generalized convergence theorem and characterization of local optimality[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,1984, PAMI-6(1): 81-87.
[11]
HEX S, HEF, FANY P,et al.An effective clustering scheme for high-dimensional data[J].Multimedia Tools and Applications,2024, 83:45001-45045.
[12]
IKOTUNA M, EZUGWUA E, ABUALIGAHL,et al.K-means clustering algorithms:a comprehensive review, variants analysis, and advances in the era of big data[J].Information Sciences,2023,622: 178-210.
HEXuansen, HEFan, YUHailan.Initialization improvement and clustering quality evaluation of K-means algorithm[J].Journal of Xi’an Polytechnic University,2024,38(6):114-123.
HEXuansen, HEFan, XULi,et al.Determination of the optimal number of clusters in K-means algorithm[J].Journal of University of Electronic Science and Technology of China,2022,51(6):904-912.
HEFan, HEXuansen, LIURunzong,et al.Data dimensionality reduction and clustering quality evaluation of K-means clustering[J].Journal of Chongqing University of Technology (Natural Science),2024,38(1):131-141.
[21]
IWASAKIY, SASAKIY, NAGATAT,et al.Dynamic mode decomposition based on expectation-maximization algorithm for simultaneous system identification and denoising[J].Mechanical Systems and Signal Processing,2025,223:111864.
[22]
XIEJ R, HONGT, LAINGT,et al.On normality assumption in residual simulation for probabilistic load forecasting[J].IEEE Transactions on Smart Grid,2017,8(3):1046-1053.
[23]
CREIGHTONJ H C.A first course in probability models and statistical inference[M].New York,NY:Springer New York,1994:140-156.
[24]
HUBERTL, ARABIEP.Comparing partitions[J].Journal of Classification,1985,2:193-218.
[25]
STEINLEYD.Properties of the hubert-arabie adjusted rand index[J]. Psychological Methods,2004,9(3):386-396.
[26]
STREHLA, GHOSHJ.Cluster ensembles-a knowledge reuse framework for combining multiple partitions[J].Journal of Machine Learning Research,2002,3:583-617.
[27]
ROSENBERGA, HIRSCHBERGJ.V-measure:a conditional entropy-based external cluster evaluation measure[C]//Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL).Prague,Czech RepublicAssociation for Computational Linguistics,2007:410-420.
[28]
FOWLKESE B, MALLOWSC L.A method for comparing two hierarchical clusterings[J].Journal of the American Statistical Association,1983,78(383):553-569.
[29]
BAGIROVA M, ALIGULIYEVR M, SULTANOVAN.Finding compact and well-separated clusters:clustering using silhouette coefficients[J].Pattern Recognition,2023,135:109144.