Skip to main navigation Skip to search Skip to main content

Morpheme Matching Based Text Tokenization for a Scarce Resourced Language

  • Zobia Rehman
  • , Waqas Anwar
  • , Usama Ijaz Bajwa
  • , Wang Xuan
  • , Zhou Chaoying
  • COMSATS University Islamabad
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalArticlepeer-review

Abstract

Text tokenization is a fundamental pre-processing step for almost all the information processing applications. This task is nontrivial for the scarce resourced languages such as Urdu, as there is inconsistent use of space between words. In this paper a morpheme matching based approach has been proposed for Urdu text tokenization, along with some other algorithms to solve the additional issues of boundary detection of compound words, affixation, reduplication, names and abbreviations. This study resulted into 97.28% precision, 93.71% recall, and 95.46% F1-measure; while tokenizing a corpus of 57000 words by using a morpheme list with 6400 entries.

Original languageEnglish
Article numbere68178
JournalPLOS ONE
Volume8
Issue number8
DOIs
StatePublished - 21 Aug 2013
Externally publishedYes

Fingerprint

Dive into the research topics of 'Morpheme Matching Based Text Tokenization for a Scarce Resourced Language'. Together they form a unique fingerprint.

Cite this