Skip to main navigation Skip to search Skip to main content

A new sampling approach for classification of imbalanced data sets with high density

  • Harbin Institute of Technology Shenzhen

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Class imbalance of datasets is a common problem in the field of machine learning. In recent years, because the traditional classifier algorithms are designed only for balanced cases, these classifiers always achieved poor performance in imbalanced data classification issues, especially for the imbalanced data with a really high density. This paper introduces the importance of imbalanced data classification in various fields first; then, contends existing methods of solving the imbalanced data classification problem; finally, proposes two new sampling methods, which are based on borderline-SMOTE, for the imbalanced data with high density, especially for big data with this kind of distribution feature. These two new algorithms are not only over-sampling the minority samples near the borderline, but also creating appropriate synthetic samples in the majority class samples side and under-sampling some particular majority class samples. Experiments show that these two algorithms could achieve a better performance than random over sampling, SMOTE (Synthetic minority over-sampling technique) and Borderline-SMOTE in AUC (Area under Receiver Operating Characteristics Curve) metric evaluate method, when the sampling rate makes the majority class and minority class samples approximate equilibrium.

Original languageEnglish
Title of host publication2014 International Conference on Big Data and Smart Computing, BIGCOMP 2014
PublisherIEEE Computer Society
Pages217-222
Number of pages6
ISBN (Print)9781479939190
DOIs
StatePublished - 2014
Externally publishedYes
Event2014 International Conference on Big Data and Smart Computing, BIGCOMP 2014 - Bangkok, Thailand
Duration: 15 Jan 201417 Jan 2014

Publication series

Name2014 International Conference on Big Data and Smart Computing, BIGCOMP 2014

Conference

Conference2014 International Conference on Big Data and Smart Computing, BIGCOMP 2014
Country/TerritoryThailand
CityBangkok
Period15/01/1417/01/14

Keywords

  • big data
  • classification
  • high density
  • imbalanced data
  • sampling method

Fingerprint

Dive into the research topics of 'A new sampling approach for classification of imbalanced data sets with high density'. Together they form a unique fingerprint.

Cite this