Skip to main navigation Skip to search Skip to main content

Cleanix: A big data cleaning parfait

  • Hongzhi Wang
  • , Mingda Li
  • , Yingyi Bu
  • , Jianzhong Li
  • , Hong Gao
  • , Jiacheng Zhang
  • Harbin Institute of Technology
  • University of California at Irvine

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In this demo, we present Cleanix, a prototype system for cleaning relational Big Data. Cleanix takes data integrated from multiple data sources and cleans them on a shared-nothing machine cluster. The backend system is built on-top-of an extensible and flexible data-parallel substrate-the Hyracks framework. Cleanix supports various data cleaning tasks such as abnormal value detection and correction, incomplete data filling, de-duplication, and conflict resolution. We demonstrate that Cleanix is a practical tool that supports effective and efficient data cleaning at the large scale.

Original languageEnglish
Title of host publicationCIKM 2014 - Proceedings of the 2014 ACM International Conference on Information and Knowledge Management
PublisherAssociation for Computing Machinery
Pages2024-2026
Number of pages3
ISBN (Electronic)9781450325981
DOIs
StatePublished - 3 Nov 2014
Event23rd ACM International Conference on Information and Knowledge Management, CIKM 2014 - Shanghai, China
Duration: 3 Nov 20147 Nov 2014

Publication series

NameCIKM 2014 - Proceedings of the 2014 ACM International Conference on Information and Knowledge Management

Conference

Conference23rd ACM International Conference on Information and Knowledge Management, CIKM 2014
Country/TerritoryChina
CityShanghai
Period3/11/147/11/14

Fingerprint

Dive into the research topics of 'Cleanix: A big data cleaning parfait'. Together they form a unique fingerprint.

Cite this