Skip to main navigation Skip to search Skip to main content

Simple semi-supervised learning for chinese word segmentation and pos tagging

  • Xinxin Li*
  • , Xuan Wang
  • , Muhammad Waqas
  • , Anwar Harbin
  • *Corresponding author for this work
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Strategies of unlabeled data selection are important for semi-supervised learning of natural language processing tasks. To increase the accuracy and diversity of new labeled data, plenty of methods have been proposed, such as ensemble-based self-training, co-training and tri-training methods. In this paper, we propose a simple and effective semi-supervised algorithm for Chinese word segmentation and part-of-speech tagging problem which selects new labeled data agreed by two different approaches: character-based and word-based models. Theoretical and experimental analysis verifies that sentences with same annotation on both models are more accurate than those generated by single models and are suitable for semi-supervised learning as additional data. Experimental results on Chinese Treebank 5.0 demonstrate that our semi-supervised approach is comparable with the best reported semi-supervised approach which employs complex feature engineering.

Original languageEnglish
Pages (from-to)5955-5961
Number of pages7
JournalInformation Technology Journal
Volume12
Issue number20
DOIs
StatePublished - 2013

Keywords

  • Chinese word segmentation and POS tagging
  • Joint decoding
  • New agreed data
  • Semi-supervised learning

Fingerprint

Dive into the research topics of 'Simple semi-supervised learning for chinese word segmentation and pos tagging'. Together they form a unique fingerprint.

Cite this