TY - GEN
T1 - A new sampling approach for classification of imbalanced data sets with high density
AU - Jia, Pengfei
AU - Zhang, Chunkai
AU - He, Zhenyu
PY - 2014
Y1 - 2014
N2 - Class imbalance of datasets is a common problem in the field of machine learning. In recent years, because the traditional classifier algorithms are designed only for balanced cases, these classifiers always achieved poor performance in imbalanced data classification issues, especially for the imbalanced data with a really high density. This paper introduces the importance of imbalanced data classification in various fields first; then, contends existing methods of solving the imbalanced data classification problem; finally, proposes two new sampling methods, which are based on borderline-SMOTE, for the imbalanced data with high density, especially for big data with this kind of distribution feature. These two new algorithms are not only over-sampling the minority samples near the borderline, but also creating appropriate synthetic samples in the majority class samples side and under-sampling some particular majority class samples. Experiments show that these two algorithms could achieve a better performance than random over sampling, SMOTE (Synthetic minority over-sampling technique) and Borderline-SMOTE in AUC (Area under Receiver Operating Characteristics Curve) metric evaluate method, when the sampling rate makes the majority class and minority class samples approximate equilibrium.
AB - Class imbalance of datasets is a common problem in the field of machine learning. In recent years, because the traditional classifier algorithms are designed only for balanced cases, these classifiers always achieved poor performance in imbalanced data classification issues, especially for the imbalanced data with a really high density. This paper introduces the importance of imbalanced data classification in various fields first; then, contends existing methods of solving the imbalanced data classification problem; finally, proposes two new sampling methods, which are based on borderline-SMOTE, for the imbalanced data with high density, especially for big data with this kind of distribution feature. These two new algorithms are not only over-sampling the minority samples near the borderline, but also creating appropriate synthetic samples in the majority class samples side and under-sampling some particular majority class samples. Experiments show that these two algorithms could achieve a better performance than random over sampling, SMOTE (Synthetic minority over-sampling technique) and Borderline-SMOTE in AUC (Area under Receiver Operating Characteristics Curve) metric evaluate method, when the sampling rate makes the majority class and minority class samples approximate equilibrium.
KW - big data
KW - classification
KW - high density
KW - imbalanced data
KW - sampling method
UR - https://www.scopus.com/pages/publications/84900662107
U2 - 10.1109/BIGCOMP.2014.6741439
DO - 10.1109/BIGCOMP.2014.6741439
M3 - 会议稿件
AN - SCOPUS:84900662107
SN - 9781479939190
T3 - 2014 International Conference on Big Data and Smart Computing, BIGCOMP 2014
SP - 217
EP - 222
BT - 2014 International Conference on Big Data and Smart Computing, BIGCOMP 2014
PB - IEEE Computer Society
T2 - 2014 International Conference on Big Data and Smart Computing, BIGCOMP 2014
Y2 - 15 January 2014 through 17 January 2014
ER -