Skip to main navigation Skip to search Skip to main content

Scalable classification by clustering: Hybrid can be better than Pure

  • Shengchun Deng*
  • , Zengyou He
  • , Xiaofei Xu
  • *Corresponding author for this work
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

The problem of scalable classification by clustering in large databases was discussed. Clustering based classification method first generates clusters using clustering algorithms. To classify new coming data points, it finds the k nearest clusters of the data point as neighbors, and assign each data point to the dominant class of these neighbors. Existing algorithms incorporated class information in making clustering decisions and produced pure clusters (each cluster associated with only one class). We presented hybrid cluster based algorithms, which produce clusters by unsupervised clustering and allow each cluster associated with multiple classes. Experimental results show that hybrid cluster based algorithms outperform pure ones in both classification accuracy and training speed.

Original languageEnglish
Pages (from-to)131-135
Number of pages5
JournalHigh Technology Letters
Volume13
Issue number2
StatePublished - Jun 2007

Keywords

  • Classification
  • Clustering
  • Data mining

Fingerprint

Dive into the research topics of 'Scalable classification by clustering: Hybrid can be better than Pure'. Together they form a unique fingerprint.

Cite this