TY - GEN
T1 - Clustering spatial data by the neighbors intersection and the density difference
AU - Yan, Zhenglong
AU - Luo, Wenjian
AU - Bu, Chenyang
AU - Ni, Li
N1 - Publisher Copyright:
© 2016 ACM.
PY - 2016/12/6
Y1 - 2016/12/6
N2 - Clustering is a classical unsupervised learning task, which is aimed to divide a data set into several groups with similar objects. Clustering problem has been studied for many years, and many excellent clustering algorithms have been proposed. In this paper, we propose a novel clustering method based on density, which is simple but effective. The primary idea of the proposed method is given as follows. Firstly, the point with the largest local density in a cluster is considered as the cluster center. The local density of each point is estimated based on the distance (called radius) between the point and its k-th nearest neighbor. The point with a smaller radius indicates a larger local density. Secondly, the difference of the local densities between each two internal points should be small, while the difference between the density of a border point and the density of an internal point should be relatively large. Thirdly, if the intersection of k nearest neighbors of two points is small, they should be assigned to different clusters. The proposed algorithm has been compared with a typical clustering algorithm named FDPCluster, and the experimental results show that our algorithm has better clustering quality.
AB - Clustering is a classical unsupervised learning task, which is aimed to divide a data set into several groups with similar objects. Clustering problem has been studied for many years, and many excellent clustering algorithms have been proposed. In this paper, we propose a novel clustering method based on density, which is simple but effective. The primary idea of the proposed method is given as follows. Firstly, the point with the largest local density in a cluster is considered as the cluster center. The local density of each point is estimated based on the distance (called radius) between the point and its k-th nearest neighbor. The point with a smaller radius indicates a larger local density. Secondly, the difference of the local densities between each two internal points should be small, while the difference between the density of a border point and the density of an internal point should be relatively large. Thirdly, if the intersection of k nearest neighbors of two points is small, they should be assigned to different clusters. The proposed algorithm has been compared with a typical clustering algorithm named FDPCluster, and the experimental results show that our algorithm has better clustering quality.
KW - Clustering
KW - Data mining
KW - Density-based custering
KW - Spatial data
UR - https://www.scopus.com/pages/publications/85013168085
U2 - 10.1145/3006299.3006332
DO - 10.1145/3006299.3006332
M3 - 会议稿件
AN - SCOPUS:85013168085
T3 - Proceedings - 3rd IEEE/ACM International Conference on Big Data Computing, Applications and Technologies, BDCAT 2016
SP - 217
EP - 226
BT - Proceedings - 3rd IEEE/ACM International Conference on Big Data Computing, Applications and Technologies, BDCAT 2016
PB - Association for Computing Machinery, Inc
T2 - 3rd IEEE/ACM International Conference on Big Data Computing, Applications and Technologies, BDCAT 2016
Y2 - 6 December 2016 through 9 December 2016
ER -