Skip to main navigation Skip to search Skip to main content

Research on Chinese lexical analysis system by fusing multiple knowledge sources

  • Wei Jiang*
  • , Xiao Long Wang
  • , Yi Guan
  • , Jian Zhao
  • *Corresponding author for this work
  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Chinese lexical analysis is the foundation task for most Chinese natural language processing. In this paper, word segmentation, POS tagging, named entity recognition and their relation are well discussed. Moreover, a pragmatic lexical analysis system based on mixed language models is presented, which adopts many models, such as n-gram, hidden Markov model, maximum entropy model, support vector machine and conditional random fields, they have good performance in the special sub-tasks. The Word Segmenter participated in the Second International Chinese Word Segmentation Bakeoff in 2005, and achieved 97.2% and 96.7% in terms of F-measure in MSR and PKU open test respectively. While the POS tagging and named entity recognition modules achieved 96.1% in precision and 88.6% in F-measure respectively in open test with the corpus that came from six-month corpora of Chinese Peoples' Daily.

Original languageEnglish
Pages (from-to)137-145
Number of pages9
JournalJisuanji Xuebao/Chinese Journal of Computers
Volume30
Issue number1
StatePublished - Jan 2007
Externally publishedYes

Keywords

  • Chinese word segmentation
  • Language model
  • Lexical analysis
  • Named entity recognition
  • Part-of-speech tagging

Fingerprint

Dive into the research topics of 'Research on Chinese lexical analysis system by fusing multiple knowledge sources'. Together they form a unique fingerprint.

Cite this