一种基于自适应加权最短路径的聚类数据集优化方法

廖志远, 顾磊

小型微型计算机系统 ›› 2026, Vol. 47 ›› Issue (9) : 2141 -2150.

小型微型计算机系统 ›› 2026, Vol. 47 ›› Issue (9) : 2141 -2150. DOI: 10.20009/j.cnki.21-1106/TP.2025-0379
算法理论与人工智能

一种基于自适应加权最短路径的聚类数据集优化方法

    廖志远, 顾磊
作者信息 +

Clustering Dataset Optimization Method Based on Adaptive Weighted Shortest Path

    LIAO Zhiyuan, GU Lei
Author information +
文章历史 +

摘要

针对传统聚类算法在高维数据处理中因特征等权重假设而导致的性能退化问题,本文提出了一种基于自适应加权最短路径的聚类数据集优化方法(HIACSP-WF).该算法借鉴了利用最短路径距离来构建稳健邻居关系并引导边界对象优化这一思想,并且进一步引入了一套数据驱动、计算式、全局一致的自适应特征加权机制.通过构建形式化的自适应加权模型,HIACSP-WF能够自动评估每个特征维度的重要程度,并生成全局特征权重向量,用于升级标准欧氏距离为加权欧氏距离.实验结果表明,HIACSP-WF在多种经典聚类算法(K-means、Agglomerative和DPC)上均展现出显著且稳定的性能提升,特别是在高维复杂数据集上表现尤为突出.与基线算法相比,HIACSP-WF的核心改进在于其自适应特征加权机制,能够有效区分重要特征和噪声特征,放大关键特征的距离差异,抑制噪声干扰,从而显著提升聚类效果.此外,HIACSP-WF还与近年来先进的聚类算法(Bombing、MDMSC和DPC-DVND)进行了对比,结果显示,结合HIACSP-WF优化后的数据集,即使使用相对简单的传统DPC算法,也能超越这些先进算法的性能,证明了数据集优化方法的优越性.

Abstract

To address the performance degradation of traditional clustering algorithms in high-dimensional data processing caused by the equal-weight assumption of features,this paper proposes a clustering dataset optimization method based on adaptive weighted shortest path (HIACSP-WF).The algorithm adopts the concept of utilizing shortest path distance to construct robust neighbor relationships and guide the optimization of boundary objects.Furthermore,it introduces a data-driven,computational,and globally consistent adaptive feature weighting mechanism.By constructing a formalized adaptive weighting model,HIACSP-WF automatically evaluates the importance of each feature dimension and generates a global feature weight vector to upgrade the standard Euclidean distance to a weighted Euclidean distance.Experimental results show that HIACSP-WF achieves significant and stable performance improvements across various classical clustering algorithms (including K-means、Agglomerative and DPC),particularly on high-dimensional complex datasets.Compared with baseline algorithms,HIACSP-WF′s core advancement lies in its adaptive feature weighting mechanism,which effectively distinguishes important features from noise features,amplifies distance differences for key features,and suppresses noise interference,thereby substantially enhancing clustering outcomes.Additionally,HIACSP-WF is evaluated against recent advanced clustering algorithms (including Bombing,MDMSC,and DPC-DVND).The results demonstrate that even a relatively simple traditional DPC algorithm,when applied to the HIACSP-WF-optimized dataset,surpasses the performance of these advanced methods,confirming the superiority of the dataset refinement approach.

关键词

聚类分析 / 特征加权 / 最短路径距离 / 自适应权重 / 数据挖掘

Key words

clustering analysis / feature weighting / shortest path distance / adaptive weighting / data mining

引用本文

引用格式 ▾
廖志远, 顾磊. 一种基于自适应加权最短路径的聚类数据集优化方法[J]. 小型微型计算机系统, 2026, 47(9): 2141-2150 DOI:10.20009/j.cnki.21-1106/TP.2025-0379

登录浏览全文

4963

注册一个新账户 忘记密码

参考文献

[1] Jain K A.Data clustering:50 years beyond K-means[J].Pattern Recognition Letters,2010,31(8):651-666.
[2] Alizadeh A A,Eisen M B,Davis R E,et al.Distinct types of diffuse large B-cell lymphoma identified by gene expression profiling[J].Nature,2000,403(6769):503-511.
[3] Yang J,Lin C T.Multi-view adjacency-constrained hierarchical clustering[J].IEEE Transactions on Emerging Topics in Computational Intelligence,2023,7(4):1126-1138.
[4] Tang J,Chang Y,Aggarwal C,et al.A survey of signed network mining in social media[J].ACM Computing Surveys,2016,49(3):1-37.
[5] Wedel M,Kamakura A W.Market segmentation:conceptual and methodological foundations[M].Berlin:Springer Science & Business Media,2000.
[6] Macqueen J B.Some methods for classification and analysis of multivariate observations[C]//5th Berkeley Symposium on Mathematical Statistics and Probability,1967:281-297.
[7] Frey B J,Dueck D.Clustering by passing messages between data points[J].Science,2007,315(5814):972-976.
[8] Alex R,Alessandro L.Clustering by fast search and find of density peaks[J].Science,2014,344(6191):1492-1496.
[9] Li Q,Wang S,Zhao C,et al.HIBOG:improving the clustering accuracy by ameliorating dataset with gravitation[J].Information Sciences,2021,550:41-56,doi:10.1016/j.ins.2020.10.046.
[10] Li Q,Wang S,Zhao C,et al.How to improve the accuracy of clustering algorithms[J].Information Sciences,2023,627:52-70,doi:10.1016/j.ins.2023.01.094.
[11] Zeng X,Wang S,Li Q,et al.Highly improve the accuracy of clustering algorithms based on shortest path distance[J].Information Sciences,2025,710:122087,doi:10.1016/J.INS.2025.122087.
[12] Aghajanyan A.Gravitational clustering[EB/OL].https://arxiv.org/abs/1509.01659.
[13] Blekas K,Lagaris I.Newtonian clustering:an approach based on molecular dynamics and global optimization[J].Pattern Recognition,2006,40(6):1734-1744.
[14] Wong K,Peng C,Li Y,et al.Herd clustering:a synergistic data clustering approach using collective intelligence[J].Applied Soft Computing Journal,2014,23:61-75,doi:10.1016/j.asoc.2014.05.034.
[15] Shi Y,Song Y,Zhang A.A shrinking-based clustering approach for multidimensional data[J].IEEE Transactions on Knowledge and Data Engineering,2005,17(10):1389-1403.
[16] Zang W,Che J,Ma L,et al.Density peaks clustering based on density voting and neighborhood diffusion[J].Information Sciences,2024,681:121209.
[17] Xu Z,Long Z,Meng H.Clustering by mining density distributions and splitting manifold structure[C]//AAAI Conference on Artificial Intelligence,2025:21842-21849.
[18] Fei Z,Zhai H,Yang J,et al.Discovering generalized clusters with adaptive mixture density-based clustering[J].Knowledge-Based Systems,2025,314:113250,doi:10.1016/J.KNOSYS.2025.113250.
[19] Vinh X N,Epps J,Bailey J.Information theoretic measures for clusterings comparison:variants,properties,normalization and correction for chance[J].Journal of Machine Learning Research,2010,11:2837-2854,doi:10.5555/1756006.1953024.
[20] Fowlkes B E,Mallows L C.A method for comparing two hierarchical clusterings[J].Journal of the American Statistical Association,2012,78(383):553-569.
[21] Rendon E,Abundez I,Arizmendi A,et al.Internal versus external cluster validation indexes[J].International Journal of Computers,Communications & Control,2011,5(1):27-34.
[22] Rosenberg A,Hirschberg J.V-measure:a conditional entropy-based external cluster evaluation measure[C]//Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning,2007:410-420.

基金资助

国家自然科学基金项目(62472232)资助.

AI Summary AI Mindmap

0

访问

0

被引

详细

导航
相关文章

AI思维导图

/