Skip to main navigation Skip to search Skip to main content

Semi-supervised learning for word sense disambiguation using parallel corpora

  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The Application of word sense disambiguation (WSD) methods based on supervised machine learning are limited by the difficulties in defining sense tags and acquiring labeled data for training. In this paper, the two problems of WSD are solved in a semi-supervised learning framework with the help of parallel corpora. The sense tags are defined automatically according to the results of word alignment on the parallel corpora. And label propagation, a graph-based semi-supervised algorithm, is employed. The experiments show that our method achieves great improvement on Chinese WSD tasks and the performances get significant growth when the scale of monolingual sentences is increasing.

Original languageEnglish
Title of host publicationProceedings - 2011 8th International Conference on Fuzzy Systems and Knowledge Discovery, FSKD 2011
Pages1490-1494
Number of pages5
DOIs
StatePublished - 2011
Event2011 8th International Conference on Fuzzy Systems and Knowledge Discovery, FSKD 2011, Jointly with the 2011 7th International Conference on Natural Computation, ICNC'11 - Shanghai, China
Duration: 26 Jul 201128 Jul 2011

Publication series

NameProceedings - 2011 8th International Conference on Fuzzy Systems and Knowledge Discovery, FSKD 2011
Volume3

Conference

Conference2011 8th International Conference on Fuzzy Systems and Knowledge Discovery, FSKD 2011, Jointly with the 2011 7th International Conference on Natural Computation, ICNC'11
Country/TerritoryChina
CityShanghai
Period26/07/1128/07/11

Keywords

  • label propagation
  • parallel corpora
  • semi-supervised learning
  • word sense disambiguation

Fingerprint

Dive into the research topics of 'Semi-supervised learning for word sense disambiguation using parallel corpora'. Together they form a unique fingerprint.

Cite this