Abstract
Strategies of unlabeled data selection are important for semi-supervised learning of natural language processing tasks. To increase the accuracy and diversity of new labeled data, plenty of methods have been proposed, such as ensemble-based self-training, co-training and tri-training methods. In this paper, we propose a simple and effective semi-supervised algorithm for Chinese word segmentation and part-of-speech tagging problem which selects new labeled data agreed by two different approaches: character-based and word-based models. Theoretical and experimental analysis verifies that sentences with same annotation on both models are more accurate than those generated by single models and are suitable for semi-supervised learning as additional data. Experimental results on Chinese Treebank 5.0 demonstrate that our semi-supervised approach is comparable with the best reported semi-supervised approach which employs complex feature engineering.
| Original language | English |
|---|---|
| Pages (from-to) | 5955-5961 |
| Number of pages | 7 |
| Journal | Information Technology Journal |
| Volume | 12 |
| Issue number | 20 |
| DOIs | |
| State | Published - 2013 |
Keywords
- Chinese word segmentation and POS tagging
- Joint decoding
- New agreed data
- Semi-supervised learning
Fingerprint
Dive into the research topics of 'Simple semi-supervised learning for chinese word segmentation and pos tagging'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver