Skip to main navigation Skip to search Skip to main content

A clustering retrieval system of chinese information

  • Xin Guang Sha*
  • , Yuan Chao Liu
  • , Ming Liu
  • , Xiao Long Wang
  • *Corresponding author for this work
  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

With tremendous and ever-growing amounts of electronic documents from World Wide Web and digital libraries, it becomes more and more difficult to get information that people really want. In order to predigest search process, people use clustering method to browse through search results. However traditional Chinese information clustering techniques are inadequate since they don't generate clusters with highly readable themes. This paper reformats the clustering problem as a salient phrase ranking problem. Given a query and its related ranked list of documents (typically a list of titles and snippets) returned from a certain Web search engine, this method first extracts and ranks salient phrases as candidate cluster theme, based on regression model of SVR (Support Vector Regression) learned from human labeled training data. The documents are assigned to relevant salient phrases to form candidate clusters, and the final clusters are generated by merging these candidate clusters. This paper also searches for a reasonable format to display the final themes of clusters, in order to help users to find the interesting documents easily. Experiment results verified our method feasible and effective.

Original languageEnglish
Title of host publication2008 International Conference on Natural Language Processing and Knowledge Engineering, NLP-KE 2008
DOIs
StatePublished - 2008
Event2008 International Conference on Natural Language Processing and Knowledge Engineering, NLP-KE 2008 - Beijing, China
Duration: 19 Oct 200822 Oct 2008

Publication series

Name2008 International Conference on Natural Language Processing and Knowledge Engineering, NLP-KE 2008

Conference

Conference2008 International Conference on Natural Language Processing and Knowledge Engineering, NLP-KE 2008
Country/TerritoryChina
CityBeijing
Period19/10/0822/10/08

Keywords

  • Document clustering
  • Performance of clustering theme
  • Salient phrase

Fingerprint

Dive into the research topics of 'A clustering retrieval system of chinese information'. Together they form a unique fingerprint.

Cite this