Skip to main navigation Skip to search Skip to main content

PengYuan@PKU: Extracting infrequent sense instance with the same N-gram pattern for the SemEval-2010 task 15

  • Peking University
  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

This paper describes our infrequent sense identification system participating in the SemEval-2010 task 15 on Infrequent Sense Identification for Mandarin Text to Speech Systems. The core system is a supervised system based on the ensembles of Naïve Bayesian classifiers. In order to solve the problem of unbalanced sense distribution, we intentionally extract only instances of infrequent sense with the same N-gram pattern as the complement training data from an untagged Chinese corpus - People's Daily of the year 2001. At the same time, we adjusted the prior probability to adapt to the distribution of the test data and tuned the smoothness coefficient to take the data sparseness into account. Official result shows that, our system ranked the first with the best Macro Accuracy 0.952. We briefly describe this system, its configuration options and the features used for this task and present some discussion of the results.

Original languageEnglish
Title of host publicationACL 2010 - SemEval 2010 - 5th International Workshop on Semantic Evaluation, Proceedings
PublisherAssociation for Computational Linguistics (ACL)
Pages371-374
Number of pages4
ISBN (Electronic)1932432701, 9781932432701
StatePublished - 2010
Event5th International Workshop on Semantic Evaluation, SemEval 2010 - Uppsala, Sweden
Duration: 15 Jul 201016 Jul 2010

Publication series

NameACL 2010 - SemEval 2010 - 5th International Workshop on Semantic Evaluation, Proceedings

Conference

Conference5th International Workshop on Semantic Evaluation, SemEval 2010
Country/TerritorySweden
CityUppsala
Period15/07/1016/07/10

Fingerprint

Dive into the research topics of 'PengYuan@PKU: Extracting infrequent sense instance with the same N-gram pattern for the SemEval-2010 task 15'. Together they form a unique fingerprint.

Cite this