Skip to main navigation Skip to search Skip to main content

Chinese text chunking based on improved K-means clustering

  • Ying Hong Liang*
  • , Tie Jun Zhao
  • , Hao Yu
  • , Jian Min Yao
  • , Bing Xu
  • *Corresponding author for this work
  • School of Computer Science and Technology, Harbin Institute of Technology
  • School of Information and Computer Engineering

Research output: Contribution to journalArticlepeer-review

Abstract

An improved k-means clustering method is proposed to identify Chinese phrases with the purpose of avoiding data sparseness and taking think of the relationship of neighbor part of speech and the cohesion of all part of speeches within one phrase. The proposed method regards each phrase as a cluster whose kernel is headword, which richly used the constituent disciplinarian of one phrase. It also integrates supervised statistical method and unsupervised clustering method by setting the original center of each class according the data from small Chinese corpus, which not only improves the accuracy of clustering but also avoids data sparseness. Through testing on Chinese Penn Treebank, the F score of seven types of Chinese phrase achieves to 92.94%. So, it is effective for Chinese text chunking.

Original languageEnglish
Pages (from-to)1106-1109
Number of pages4
JournalHarbin Gongye Daxue Xuebao/Journal of Harbin Institute of Technology
Volume39
Issue number7
StatePublished - Jul 2007
Externally publishedYes

Keywords

  • Chinese text chunking
  • K-means clustering
  • Sparseness

Fingerprint

Dive into the research topics of 'Chinese text chunking based on improved K-means clustering'. Together they form a unique fingerprint.

Cite this