基于知识诱导与无用中心驱动的K-means算法
DOI:
作者:
作者单位:

华东交通大学理学院,江西南昌 330013

作者简介:

王森(1969—),男,教授,硕士生导师,研究方向为计算机算法与应用。E-mail:wangsen@ecjtu.edu.cn。

通讯作者:

中图分类号:

TP391

基金项目:

国家自然科学基金项目(12361004)


K-means Algorithm Driven by Knowledge Induction and Useless Center
Author:
Affiliation:

School of Science, East China Jiaotong University, Nanchang 330013 , China

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    K-means算法是一种广泛应用的高效无监督聚类算法。然而,研究表明,在处理高维或非球状结构的数据集时,K-means算法在确定聚类数目和选择初始质心方面存在显著局限性。为优化K-means算法的初始质心选取机制,解决聚类数目确定问题,文章提出了一种基于知识诱导与无用中心驱动的K-means算法。该算法首先引入高密度知识点的检测机制,通过识别数据集中的高密度知识点,构建候选质心集合,基于高斯混合模型理论推导最优聚类数;随后,采用无用中心筛选策略对候选质心进行优化选择,最终确定最优初始质心集合。在真实数据集上的实验结果表明,所提算法在聚类性能上总体优于其他对比算法。该算法可有效解决非球状数据分布的聚类问题,并在复杂数据结构场景下展现出较为优越的聚类性能。

    Abstract:

    K-means is a widely used and efficient unsupervised clustering algorithm. However, studies have shown that when dealing with high-dimensional or non-spherically distributed datasets, the K-means algorithm has significant limitations in determining the number of clusters and selecting initial centroids. To thoroughly explore and optimize the initial centroid selection mechanism and the problem of determining the number of clusters in the K-means algorithm, a K-means algorithm based on knowledge induction and useless center driven is proposed. This algorithm first introduces a detection mechanism for high-density knowledge points, constructs a candidate centroid set by identifying high-density knowledge points in the dataset; then infers the optimal number of clusters based on Gaussian mixture model theory; subsequently adopts a useless center screening strategy to optimally select the candidate centroids, and finally determines the optimal initial centroid set. Experiments on real datasets show that the proposed optimized algorithm generally outperforms other comparison algorithms in clustering performance. This algorithm effectively solves the clustering problem of non-spherical data distribution and exhibits relatively superior clustering performance in scenarios with complex data structures.

    参考文献
    相似文献
    引证文献
引用本文

王森,刘青阳,詹小秦,等. 基于知识诱导与无用中心驱动的K-means算法[J]. 华东交通大学学报,2026,43(3):120-126.

复制
分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-03-25
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-07-16
  • 出版日期:
关闭