Abstract
Many practical applications suffer from imbalanced data classification, in which case the minority class has degraded recognition rate. The primary causes are the sample scarcity of the minority class and the intrinsic complex distribution characteristics of imbalanced datasets. The imbalanced classification problem is more serious on small sample datasets. To solve the problems of small sample and class imbalance, a hybrid resampling method is proposed. The proposed method combines an oversampling approach (synthetic minority oversampling technique, SMOTE) and a novel data cleaning approach (weighted edited nearest neighbor rule, WENN). First, SMOTE generates synthetic minority class examples using linear interpolation. Then, WENN detects and deletes unsafe majority and minority class examples using weighted distance function and k-nearest neighbor (kNN) rule. The weighted distance function scales up a commonly used distance by considering local imbalance and spacial sparsity. Extensive experiments over synthetic and real datasets validate the superiority of the proposed SMOTE-WENN compared with three state-of-the-art resampling methods.
| Original language | English |
|---|---|
| Pages (from-to) | 1394-1409 |
| Number of pages | 16 |
| Journal | Applied Intelligence |
| Volume | 51 |
| Issue number | 3 |
| DOIs | |
| State | Published - Mar 2021 |
| Externally published | Yes |
Keywords
- Data cleaning
- Imbalanced data classification
- Oversampling
- Small sample datasets
Fingerprint
Dive into the research topics of 'SMOTE-WENN: Solving class imbalance and small sample problems by oversampling and distance scaling'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver