Abstract
This paper presents a new efficient algorithm for clustering categorical data, Squeezer, which can produce high quality clustering results and at the same time deserve good scalability. The Squeezer algorithm reads each tuple t in sequence, either assigning t to an existing cluster (initially none), or creating t as a new cluster, which is determined by the similarities between t and clusters. Due to its characteristics, the proposed algorithm is extremely suitable for clustering data streams, where given a sequence of points, the objective is to maintain consistently good clustering of the sequence so far, using a small amount of memory and time. Outliers can also be handled efficiently and directly in Squeezer. Experimental results on real-life and synthetic datasets verify the superiority of Squeezer.
| Original language | English |
|---|---|
| Pages (from-to) | 611-624 |
| Number of pages | 14 |
| Journal | Journal of Computer Science and Technology |
| Volume | 17 |
| Issue number | 5 |
| DOIs | |
| State | Published - Sep 2002 |
Keywords
- Categorical data
- Clustering
- Data mining
- Data stream
Fingerprint
Dive into the research topics of 'Squeezer: An efficient algorithm for clustering categorical data'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver