Skip to main navigation Skip to search Skip to main content

A topic clustering approach to finding similar questions from large question and answer archives

  • National University of Singapore
  • Xiamen University
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

With the blooming of Web 2.0, Community Question Answering (CQA) services such as Yahoo! Answers (http://answers.yahoo.com), WikiAnswer (http://wiki.answers.com), and Baidu Zhidao (http://zhidao.baidu.com), etc., have emerged as alternatives for knowledge and information acquisition. Over time, a large number of question and answer (Q&A) pairs with high quality devoted by human intelligence have been accumulated as a comprehensive knowledge base. Unlike the search engines, which return long lists of results, searching in the CQA services can obtain the correct answers to the question queries by automatically finding similar questions that have already been answered by other users. Hence, it greatly improves the efficiency of the online information retrieval. However, given a question query, finding the similar and wellanswered questions is a non-trivial task. The main challenge is the word mismatch between question query (query) and candidate question for retrieval (question ). To investigate this problem, in this study, we capture the word semantic similarity between query and question by introducing the topic modeling approach. We then propose an unsupervised machine-learning approach to finding similar questions on CQA Q&A archives. The experimental results show that our proposed approach significantly outperforms the state-of-the-art methods.

Original languageEnglish
Article numbere71511
JournalPLOS ONE
Volume9
Issue number3
DOIs
StatePublished - 4 Mar 2014

Fingerprint

Dive into the research topics of 'A topic clustering approach to finding similar questions from large question and answer archives'. Together they form a unique fingerprint.

Cite this