TY - GEN
T1 - A Semi-supervised Clustering Method through Bottleneck Distance Exploration
AU - Yao, Yuan
AU - Li, Yan
AU - Wang, Ke
AU - Huang, Zhichao
AU - Ye, Yunming
N1 - Publisher Copyright:
© 2016 IEEE.
PY - 2016/6/28
Y1 - 2016/6/28
N2 - Semi-supervised clustering is one of the most active research area in machine learning and pattern recognition, which can improve the performance of unsupervised clustering efficiently. This paper focuses on exploiting both the label information of a few labeled samples and the spatial distribution information of large amount of unlabeled samples. We proposed a new semi-supervised clustering method, named Bottleneck Distance based Semi-supervised Clustering (BDSC), which is based on the idea of label propagation and can perform clustering with no parameters. BDSC works by firstly obtaining small amount of labeled samples for each class. Then, a minimum spanning tree is constructed from both labeled and unlabeled samples, where the distances between an unlabeled sample and labeled samples are computed to get the bottleneck distance for each unlabeled sample. Finally, labels are propagated by comparing the bottleneck distances. Experimental results demonstrate that the proposed technique outperforms classical clustering algorithms with respect to the precision and the capability of recognizing nonspherical-shaped clusters.
AB - Semi-supervised clustering is one of the most active research area in machine learning and pattern recognition, which can improve the performance of unsupervised clustering efficiently. This paper focuses on exploiting both the label information of a few labeled samples and the spatial distribution information of large amount of unlabeled samples. We proposed a new semi-supervised clustering method, named Bottleneck Distance based Semi-supervised Clustering (BDSC), which is based on the idea of label propagation and can perform clustering with no parameters. BDSC works by firstly obtaining small amount of labeled samples for each class. Then, a minimum spanning tree is constructed from both labeled and unlabeled samples, where the distances between an unlabeled sample and labeled samples are computed to get the bottleneck distance for each unlabeled sample. Finally, labels are propagated by comparing the bottleneck distances. Experimental results demonstrate that the proposed technique outperforms classical clustering algorithms with respect to the precision and the capability of recognizing nonspherical-shaped clusters.
KW - bottleneck distance
KW - label propagation
KW - minimum spanning tree
KW - semi-supervised clustering
UR - https://www.scopus.com/pages/publications/85034235461
U2 - 10.1109/ICSS.2016.22
DO - 10.1109/ICSS.2016.22
M3 - 会议稿件
AN - SCOPUS:85034235461
T3 - Proceedings of International Conference on Service Science, ICSS
SP - 115
EP - 121
BT - Proceedings - 2016 9th International Conference on Service Science, ICSS 2016
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 9th International Conference on Service Science, ICSS 2016
Y2 - 15 October 2016 through 16 October 2016
ER -