School of Electronic and Information Engineering,Lanzhou Jiaotong University,Lanzhou 730070,China
Show less
文章历史+
Received
Published
2025-06-24
2026-06-25
Issue Date
2026-08-18
PDF (4390K)
摘要
针对模糊聚类算法在处理时间序列时未全面关注序列分布特征等问题,提出一种考虑序列分布特征的时间序列模糊聚类(considering sequence distribution characteristics time series fuzzy clustering, CDF).CDF通过三方面强化聚类过程:提出关联序列分布特征的初始簇中心确定方法,优化簇中心的选择;构建感知序列分布特征的距离度量,提升算法对噪声的鲁棒性;提出基于序列结构的自适应模糊因子策略,更新隶属度矩阵,持续优化簇中心直至收敛到最优解.CDF充分考虑序列的高维复杂分布结构,能有效处理序列的模糊性,提升聚类结果的可解释性.为验证CDF的性能,将其与8种先进的聚类算法在UCR数据库的7个时间序列数据集上进行了对比.实验结果表明,CDF在2个量化指标上优于所有对比方法,并在部分数据集上展现出高抗噪性、强收敛性及具有竞争力的计算效率.
Abstract
To address the problem that fuzzy clustering algorithms do not fully account for sequence distribution characteristics when processing time series, we propose a time series fuzzy clustering (CDF) method considering sequence distribution characteristic. CDF strengthens the clustering process through three aspects: proposing an initial cluster center determination method for the sequence distribution characteristics, optimizing the selection of cluster centers; constructing a distance metric for the distribution characteristics of perceptual sequences to enhance the robustness of algorithm to noise; proposing an adaptive fuzzy factor strategy based on sequence structures, updating the membership matrix, and continuously optimizing the cluster center until it converges to the optimal solution. CDF fully accounts for the high-dimensional complex distribution structure of sequences, enabling more effective handling of ambiguity in sequences and improving the interpretability of clustering results. To validate the performance of CDF, it was compared with 8 advanced clustering algorithms on 7 time-series datasets from the UCR database. Experimental results demonstrate that CDF outperforms all baselines across two metrics, exhibiting high noise robustness, strong convergence, and competitive computational efficiency on some datasets.
近年来,模糊C均值(fuzzy C means, FCM)由于具有软聚类优势,已成为时间序列聚类中的重要方法[4].与传统的硬聚类相比,FCM允许每个序列同时属于多个簇,通过隶属度反映其模糊性与不确定性.因此,FCM更适合描述具有模糊边界的复杂序列数据.但现有FCM改进算法在三个关键环节——初始化、距离度量和隶属度更新——都未充分考虑序列的分布特征.例如,缺乏对序列动态变化的敏感性,导致在面临复杂、异构的序列时,算法容易陷入局部最优或局部偏差.这表明有必要设计一种融入序列分布特征的聚类思想,以增强其鲁棒性和准确性.
为解决上述问题,提出一种考虑序列分布特征的时间序列模糊聚类(considering sequence distribution characteristics time series fuzzy clustering, CDF),CDF通过分布感知策略,实现全流程的闭环:(1) 初始化:提出基于序列分布特征的簇中心确定方法,分析序列的分布特征,以选择更符合数据分布的初始簇中心,提升聚类的稳定性与准确性;(2) 距离度量:定义一种感知序列分布特征的欧氏距离(sensing the euclidean distance characterizing the distribution of sequences, SED),使距离不依赖于数据的量纲,增强算法在不同数据尺度及波动情况下的一致性和鲁棒性;(3) 隶属度更新:提出基于序列结构的自适应模糊因子策略,通过调整模糊因子,使其考虑序列的分布特性,从而增强对复杂关系的捕捉能力.此外,为降低高维序列的计算复杂度,CDF在预处理阶段执行归一化操作,统一尺度、压缩数值范围,为后续阶段提供更高效、一致的输入空间.多步协同能够减小序列波动和异常值对聚类效果的影响,提升整体的稳定性和鲁棒性.
HUANGZ, HAOH, DUL .Exploring the explainability of time series clustering:a review of methods and practices[C]//Proceedings of the Eighteenth ACM International Conference on Web Search and Data Mining.Hannover Germany.ACM,2025:1005-1007.
LIH L, ZHANGL P. Summary of clustering research in time series data mining[J]. Journal of University of Electronic Science and Technology of China,2022,51(3):416-424.(in Chinese)
[4]
PAPARRIZOSJ, BOGIREDDYS P T R .Time-series clustering:a comprehensive study of data mining,machine learning,and deep learning methods[J]. Proceedings of the VLDB Endowment,2025,18(11): 4380-4395.
ZHANGC, CHENM. Density-driven time series fuzzy clustering[J/OL]. Journal of Xidian University, 2026, 53(1): 222-234. (in Chinese)
[7]
LIUZ, ZHUS J, SENAPATIT, et al .New distance measures of complex Fermatean fuzzy sets with applications in decision making and clustering problems[J].Information Sciences,2025,686:121310.
[8]
GAOY L, WANGZ H, XIEJ X, et al .A new robust fuzzy c-means clustering method based on adaptive elastic distance[J].Knowledge-Based Systems,2022,237:107769.
[9]
YUB, WUC Y. Fuzzy clustering of time series based on trend feature information granulation[J].Fuzzy Sets and Systems,2025,519:109522.
[10]
MAZ L, LÓPEZ-ORIONAÁ, OMBAOH, et al .FCPCA: fuzzy clustering of high-dimensional time series based on common principal component analysis[J]. International Journal of Approximate Reasoning,2025,187:109552.
[11]
ZHANGC B, CHENL, ZHAOY P, et al .Graph enhanced fuzzy clustering for categorical data using a Bayesian dissimilarity measure[J].IEEE Transactions on Fuzzy Systems,2023,31(3):810-824.
[12]
HASHEMZADEHM, GOLZARI OSKOUEIA, FARAJZADEHN .New fuzzy C-means clustering method based on feature-weight and cluster-weight learning[J].Applied Soft Computing,2019,78:324-345.
[13]
WUC M, ZHANGX L. A self-learning iterative weighted possibilistic fuzzy c-means clustering via adaptive fusion[J]. Expert Systems with Applications,2022,209:118280.
[14]
CHENQ, YUW Z, NIEF P, et al .Adaptive fuzzy C-means with graph embedding[EB/OL]. [2024-05-22]
[15]
WUC M, HOUJ. New semi-supervised fuzzy C-means clustering with asymmetric deviation constraints and fast algorithm[J]. Expert Systems with Applications,2026,298:129648.
[16]
GOUDAH A, AHMEDM A, ROUSHDYM I. Optimizing anomaly-based attack detection using classification machine learning[J]. Neural Computing and Applications,2024,36(6):3239-3257.
[17]
BATESS, HASTIET, TIBSHIRANIR. Cross-validation: what does it estimate and how well does it do it?[J]. Journal of the American Statistical Association,2024,119(546):1434-1445.
SEALA, KARLEKARA, KREJCARO, et al. Fuzzy c-means clustering using Jeffreys-divergence based similarity measure[J].Applied Soft Computing,2020,88:106016.
[20]
JORGEM B, RUBÉNC. Time series clustering with random convolutional kernels[J]. Data Mining and Knowledge Discovery,2024, 38(4): 1862-1888.
[21]
SVIRSKYJ, LINDENBAUMO .Interpretable deep clustering for tabular data[EB/OL].[2023-06-07]
[22]
CAIB R, HUANGG Y, YANGS Q, et al .SE-shapelets:semi-supervised clustering of time series using representative shapelets[J].Expert Systems with Applications,2024,240:122584.
[23]
PAPARRIZOSJ, GRAVANOL. K-shape:efficient and accurate clustering of time series[J]. ACM SIGMOD Record,2016,45(1):69-76.
[24]
RODRIGUEZA, LAIOA .Clustering by fast search and find of density peaks[J].Science,2014,344(6191):1492-1496.
[25]
PENGF R, LUOJ C, LUX, et al .Cross-domain contrastive learning for time series clustering[J].Proceedings of the AAAI Conference on Artificial Intelligence,2024,38(8):8921-8929.
LUH Y, FANY L, GAON, et al .An incremental density-based clustering algorithm for concept drift detection and adaption over data stream[J].Acta Electronica Sinica,2025,53(6):2050-2062.(in Chinese)
QIANL X, CHENM, MAX Y,et al .Multi-view clustering based on adaptive tensor singular value shrinkage[J].Journal of Computer Research and Development,2025,62(3):733-750.(in Chinese)
基金资助
国家自然科学基金资助项目(62266029)
National Natural ScienceFoundation of China(62266029)
甘肃省重点研发计划项目(24YFGA036)
Gansu Provincial Key Research and Development Program Projects(24YFGA036)
甘肃省联合科研基金重点项目(25JRRA1103)
Key Project of the Gansu Provincial Joint Research Fund(25JRRA1103)