Skip to main navigation Skip to search Skip to main content

SMOTE-WENN: Solving class imbalance and small sample problems by oversampling and distance scaling

  • Hongjiao Guan*
  • , Yingtao Zhang
  • , Min Xian
  • , H. D. Cheng
  • , Xianglong Tang
  • *Corresponding author for this work
  • Qilu University of Technology
  • School of Computer Science and Technology, Harbin Institute of Technology
  • University of Idaho
  • Utah State University

Research output: Contribution to journalArticlepeer-review

Abstract

Many practical applications suffer from imbalanced data classification, in which case the minority class has degraded recognition rate. The primary causes are the sample scarcity of the minority class and the intrinsic complex distribution characteristics of imbalanced datasets. The imbalanced classification problem is more serious on small sample datasets. To solve the problems of small sample and class imbalance, a hybrid resampling method is proposed. The proposed method combines an oversampling approach (synthetic minority oversampling technique, SMOTE) and a novel data cleaning approach (weighted edited nearest neighbor rule, WENN). First, SMOTE generates synthetic minority class examples using linear interpolation. Then, WENN detects and deletes unsafe majority and minority class examples using weighted distance function and k-nearest neighbor (kNN) rule. The weighted distance function scales up a commonly used distance by considering local imbalance and spacial sparsity. Extensive experiments over synthetic and real datasets validate the superiority of the proposed SMOTE-WENN compared with three state-of-the-art resampling methods.

Original languageEnglish
Pages (from-to)1394-1409
Number of pages16
JournalApplied Intelligence
Volume51
Issue number3
DOIs
StatePublished - Mar 2021
Externally publishedYes

Keywords

  • Data cleaning
  • Imbalanced data classification
  • Oversampling
  • Small sample datasets

Fingerprint

Dive into the research topics of 'SMOTE-WENN: Solving class imbalance and small sample problems by oversampling and distance scaling'. Together they form a unique fingerprint.

Cite this