Skip to main navigation Skip to search Skip to main content

Incorporating linguistic rules in statistical Chinese language model for Pinyin-to-character conversion

  • B. Liu*
  • , X. Wang
  • , Y. Wang
  • *Corresponding author for this work
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

An N-gram Chinese language model incorporating linguistic rules is presented. By constructing elements lattice, rules information is incorporated in statistical frame. To facilitate the hybrid modeling, novel methods such as MI-based rule evaluating, weighted rule quantification and element-based n-gram probability approximation are presented. Dynamic Viterbi algorithm is adopted to search the best path in lattice. To strengthen the model, transformation-based error-driven rules learning is adopted. Applying proposed model to Chinese Pinyin-to-character conversion, high performance has been achieved in accuracy, flexibility and robustness simultaneously. Tests show correct rate achieves 94.81% instead of 90.53% using bi-gram Markov model alone. Many long-distance dependency and recursion in language can be processed effectively.

Original languageEnglish
Pages (from-to)8-13
Number of pages6
JournalHigh Technology Letters
Volume7
Issue number2
StatePublished - Jun 2001

Keywords

  • Chinese Pinyin-to-character conversion
  • Element lattice
  • Hybrid language model
  • N-gram language model
  • Rule-based language model
  • Transformation-based error-driven learning

Fingerprint

Dive into the research topics of 'Incorporating linguistic rules in statistical Chinese language model for Pinyin-to-character conversion'. Together they form a unique fingerprint.

Cite this