Skip to main navigation Skip to search Skip to main content

Incomplete data classification based on multiple views

  • Ming Sun
  • , Hongzhi Wang*
  • , Fanshan Meng
  • , Jianzhong Li
  • , Hong Gao
  • *Corresponding author for this work
  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Missing values have negative impacts on big data analysis. However, in absence of extra knowledge, exact imputation can hardly be conducted for many data sets. Therefore, we have to tolerate missing values and perform data mining on incomplete data sets directly. To achieve high quality data mining on incomplete data, we propose a classification approach based on multiple views. We use various complete views of the data set to generate the base classifiers and combine the results of base classifiers. Since the amount of base classifiers will affect the effectiveness and efficiency of the classification, we aim to find proper view sets. We prove that the view set selection problem is an NP-hard problem and develop an approximation algorithm with approximate ratio ln|S| + 1 where S is the feature set of original data set. Extensive experimental results demonstrate the efficiency and effectiveness of the proposed approaches.

Original languageEnglish
Title of host publicationWeb Technologies and Applications - 18th Asia-Pacific Web Conference, APWeb 2016, Proceedings
EditorsKyuseok Shim, Kai Zheng, Guanfeng Liu, Feifei Li
PublisherSpringer Verlag
Pages239-250
Number of pages12
ISBN (Print)9783319458168
DOIs
StatePublished - 2016

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume9932 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Fingerprint

Dive into the research topics of 'Incomplete data classification based on multiple views'. Together they form a unique fingerprint.

Cite this