Abstract
A statistical method based on information entropy is proposed for domain-specific term extraction from domain comparative corpora. It takes into account the distribution of a candidate word among domains and within a certain domain. Normalization step is added into the extraction process to cope with unbalanced corpora. The proposed method characterizes attributes of domain-specific term more precisely and more effectively than previous term extraction approaches. Domain-specific terms are applied in text classification as the feature space. Experimental results indicate that it achieves better performance than traditional feature selection methods.
| Original language | English |
|---|---|
| Pages (from-to) | 328-332 |
| Number of pages | 5 |
| Journal | Tien Tzu Hsueh Pao/Acta Electronica Sinica |
| Volume | 35 |
| Issue number | 2 |
| State | Published - Feb 2007 |
| Externally published | Yes |
Keywords
- Domain-specific term
- Feature selection
- Information entropy
- Normalization
- Text classification
Fingerprint
Dive into the research topics of 'Automatic domain-specific term extraction and its application in text classification'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver