Skip to main navigation Skip to search Skip to main content

Squeezer: An efficient algorithm for clustering categorical data

  • Zengyou He*
  • , Xiaofei Xu
  • , Shengchun Deng
  • *Corresponding author for this work
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

This paper presents a new efficient algorithm for clustering categorical data, Squeezer, which can produce high quality clustering results and at the same time deserve good scalability. The Squeezer algorithm reads each tuple t in sequence, either assigning t to an existing cluster (initially none), or creating t as a new cluster, which is determined by the similarities between t and clusters. Due to its characteristics, the proposed algorithm is extremely suitable for clustering data streams, where given a sequence of points, the objective is to maintain consistently good clustering of the sequence so far, using a small amount of memory and time. Outliers can also be handled efficiently and directly in Squeezer. Experimental results on real-life and synthetic datasets verify the superiority of Squeezer.

Original languageEnglish
Pages (from-to)611-624
Number of pages14
JournalJournal of Computer Science and Technology
Volume17
Issue number5
DOIs
StatePublished - Sep 2002

Keywords

  • Categorical data
  • Clustering
  • Data mining
  • Data stream

Fingerprint

Dive into the research topics of 'Squeezer: An efficient algorithm for clustering categorical data'. Together they form a unique fingerprint.

Cite this