Skip to main navigation Skip to search Skip to main content

Automatic domain-specific term extraction and its application in text classification

  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

A statistical method based on information entropy is proposed for domain-specific term extraction from domain comparative corpora. It takes into account the distribution of a candidate word among domains and within a certain domain. Normalization step is added into the extraction process to cope with unbalanced corpora. The proposed method characterizes attributes of domain-specific term more precisely and more effectively than previous term extraction approaches. Domain-specific terms are applied in text classification as the feature space. Experimental results indicate that it achieves better performance than traditional feature selection methods.

Original languageEnglish
Pages (from-to)328-332
Number of pages5
JournalTien Tzu Hsueh Pao/Acta Electronica Sinica
Volume35
Issue number2
StatePublished - Feb 2007
Externally publishedYes

Keywords

  • Domain-specific term
  • Feature selection
  • Information entropy
  • Normalization
  • Text classification

Fingerprint

Dive into the research topics of 'Automatic domain-specific term extraction and its application in text classification'. Together they form a unique fingerprint.

Cite this