Abstract
An N-gram Chinese language model incorporating linguistic rules is presented. By constructing elements lattice, rules information is incorporated in statistical frame. To facilitate the hybrid modeling, novel methods such as MI-based rule evaluating, weighted rule quantification and element-based n-gram probability approximation are presented. Dynamic Viterbi algorithm is adopted to search the best path in lattice. To strengthen the model, transformation-based error-driven rules learning is adopted. Applying proposed model to Chinese Pinyin-to-character conversion, high performance has been achieved in accuracy, flexibility and robustness simultaneously. Tests show correct rate achieves 94.81% instead of 90.53% using bi-gram Markov model alone. Many long-distance dependency and recursion in language can be processed effectively.
| Original language | English |
|---|---|
| Pages (from-to) | 8-13 |
| Number of pages | 6 |
| Journal | High Technology Letters |
| Volume | 7 |
| Issue number | 2 |
| State | Published - Jun 2001 |
Keywords
- Chinese Pinyin-to-character conversion
- Element lattice
- Hybrid language model
- N-gram language model
- Rule-based language model
- Transformation-based error-driven learning
Fingerprint
Dive into the research topics of 'Incorporating linguistic rules in statistical Chinese language model for Pinyin-to-character conversion'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver