Skip to main navigation Skip to search Skip to main content

Template based affix stemmer for a morphologically rich language

  • Sajjad Khan
  • , Waqas Anwar
  • , Usama Bajwa
  • , Xuan Wang
  • COMSATS University Islamabad
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalArticlepeer-review

Abstract

Word stemming is one of the most significant factors that affect the performance of a Natural Language Processing (NLP) application such as Information Retrieval (IR) system, part of speech tagging, machine translation system and syntactic parsing. Urdu language raises several challenges to NLP largely due to its rich morphology. In Urdu language, stemming process is different as compared to that for other languages, as it not only depends on removing prefixes and suffixes but also on removing infixes. In this paper, we introduce a template based stemmer that eliminates all kinds of affixes i.e., prefixes, infixes and suffixes, depending on the morphological pattern of the word. The presented results are excellent and this stemmer can prove to be very affective for a morphologically rich language.

Original languageEnglish
Pages (from-to)146-154
Number of pages9
JournalInternational Arab Journal of Information Technology
Volume12
Issue number2
StatePublished - 2015
Externally publishedYes

Keywords

  • Exception lists
  • IR
  • Infix
  • Prefix
  • Stemming
  • Suffix

Fingerprint

Dive into the research topics of 'Template based affix stemmer for a morphologically rich language'. Together they form a unique fingerprint.

Cite this