Skip to main navigation Skip to search Skip to main content

Effective density-based clustering algorithms for incomplete data

  • Zhonghao Xue
  • , Hongzhi Wang*
  • *Corresponding author for this work
  • University of Southern California

Research output: Contribution to journalArticlepeer-review

Abstract

Density-based clustering is an important category among clustering algorithms. In real applications, manydatasets suffer from incompleteness. Traditional imputation technologies or other techniques for handling missingvalues are not suitable for density-based clustering and decrease clustering result quality. To avoid these problems, we develop a novel density-based clustering approach for incomplete data based on Bayesian theory, which conductsimputation and clustering concurrently and makes use of intermediate clustering results. To avoid the impact oflow-density areas inside non-convex clusters, we introduce a local imputation clustering algorithm, which aims toimpute points to high-density local areas. The performances of the proposed algorithms are evaluated using tensynthetic datasets and five real-world datasets with induced missing values. The experimental results show theeffectiveness of the proposed algorithms.

Original languageEnglish
Article number9430134
Pages (from-to)183-194
Number of pages12
JournalBig Data Mining and Analytics
Volume4
Issue number3
DOIs
StatePublished - Sep 2021

Keywords

  • clustering algorihtm
  • density-based clustering
  • incomplete data

Fingerprint

Dive into the research topics of 'Effective density-based clustering algorithms for incomplete data'. Together they form a unique fingerprint.

Cite this