Skip to main navigation Skip to search Skip to main content

Classification method for imbalance data set based on hybrid strategy

  • Peng Li*
  • , Xiao Long Wang
  • , Yuan Chao Liu
  • , Bao Xun Wang
  • *Corresponding author for this work
  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

This paper presents a novel and effective classification method for imbalanced data sets. The core idea of the algorithm, which is composed of three parts, is to provide a general solution for IDS classification by both sample preprocessing and classifier improving. Firstly, we re-sample the imbalance data by using variable SOM clustering so as to overcome the flaws of the traditional re-sampling methods, such as serious randomness, subjective interference and information loss. Then we cut down the sampled data sets according to the K-NN rule to solve the problem of data confusion, which improves the generalization of SVM. Especially, in order to adapt the class imbalance, the class boundary alignment is introduced through conformal transform on kernel function. The comparison results show the effectiveness of three algorithms. Meanwhile, the algorithm has also been used in our question answer system, which obtains outstanding result in the international TREC-2006 QA track.

Original languageEnglish
Pages (from-to)2161-2165
Number of pages5
JournalTien Tzu Hsueh Pao/Acta Electronica Sinica
Volume35
Issue number11
StatePublished - Nov 2007
Externally publishedYes

Keywords

  • Classification
  • Imbalanced data sets (IDS)
  • K-nearest neighbor (K-NN)
  • Support vector machine (SVM)
  • Variable self-organizing maps (V-SOM)

Fingerprint

Dive into the research topics of 'Classification method for imbalance data set based on hybrid strategy'. Together they form a unique fingerprint.

Cite this