Skip to main navigation Skip to search Skip to main content

Hidden Markov model based part of speech tagger for Urdu

  • Waqas Anwar*
  • , Wang Xuan
  • , LuLi
  • , Wang Xiaolong
  • *Corresponding author for this work
  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

In this study, we present the preliminary achievement of Hidden Markov Model (HMM) to solve the part of speech tagging problem of Urdu language. The presented HMM is derived from the combination of lexical and transition probabilities. An important feature of our tagger is to combine many distinguished smoothing techniques with HMM model to resolve the data sparseness problem. We note that the proposed HMM based Urdu Part of speech tagger with different smoothing method has achieved significant performance. We evaluate our tagger's results regarding different smoothing methods and different word level accuracy through Analysis of Variance (ANOVA) and show how present results are significant. Also, we compose a confusion matrix about most frequent error occurring tag pairs. The development of our tagger is an important milestone toward Urdu language processing. This will open some novel research directions to mature Urdu language processing.

Original languageEnglish
Pages (from-to)1190-1198
Number of pages9
JournalInformation Technology Journal
Volume6
Issue number8
DOIs
StatePublished - 15 Nov 2007
Externally publishedYes

Keywords

  • Hidden Markov model
  • Part of speech tagging
  • Smoothing methods
  • Urdu language

Fingerprint

Dive into the research topics of 'Hidden Markov model based part of speech tagger for Urdu'. Together they form a unique fingerprint.

Cite this