Abstract
In this study, we present the preliminary achievement of Hidden Markov Model (HMM) to solve the part of speech tagging problem of Urdu language. The presented HMM is derived from the combination of lexical and transition probabilities. An important feature of our tagger is to combine many distinguished smoothing techniques with HMM model to resolve the data sparseness problem. We note that the proposed HMM based Urdu Part of speech tagger with different smoothing method has achieved significant performance. We evaluate our tagger's results regarding different smoothing methods and different word level accuracy through Analysis of Variance (ANOVA) and show how present results are significant. Also, we compose a confusion matrix about most frequent error occurring tag pairs. The development of our tagger is an important milestone toward Urdu language processing. This will open some novel research directions to mature Urdu language processing.
| Original language | English |
|---|---|
| Pages (from-to) | 1190-1198 |
| Number of pages | 9 |
| Journal | Information Technology Journal |
| Volume | 6 |
| Issue number | 8 |
| DOIs | |
| State | Published - 15 Nov 2007 |
| Externally published | Yes |
Keywords
- Hidden Markov model
- Part of speech tagging
- Smoothing methods
- Urdu language
Fingerprint
Dive into the research topics of 'Hidden Markov model based part of speech tagger for Urdu'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver