Skip to main navigation Skip to search Skip to main content

Bilingual lexicon extraction using locally weighted linear regression from comparable corpora

  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Recently a simple linear transformation with word embedding has been found to be highly effective to extract a bilingual lexicon from comparable corpora. However, it is easy to underfit for transforming all the words just using a single transformation matrix. This paper proposes a simple non-parameter based solution using locally weighted linear regression (LWR) which forces that the closer words in the training lexicon with the target word should be more important for estimating the objective function for the regression. The experimental results confirm that the proposed solution can achieve a 36.7% relative improvement at Top-1 over the baseline approach on the English-to-Chinese bilingual lexicon extraction task.

Original languageEnglish
Title of host publicationProceedings of 2015 International Conference on Asian Language Processing, IALP 2015
EditorsBin Ma, Min Zhang, Yanfeng Lu, Minghui Dong, Wenliang Chen
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages13-16
Number of pages4
ISBN (Electronic)9781467395953
DOIs
StatePublished - 12 Apr 2016
Externally publishedYes
EventInternational Conference on Asian Language Processing, IALP 2015 - Suzhou, China
Duration: 24 Oct 201525 Oct 2015

Publication series

NameProceedings of 2015 International Conference on Asian Language Processing, IALP 2015

Conference

ConferenceInternational Conference on Asian Language Processing, IALP 2015
Country/TerritoryChina
CitySuzhou
Period24/10/1525/10/15

Keywords

  • bilingual lexicon extraction
  • locally weighted linear regression
  • transformation matrix
  • word embedding

Fingerprint

Dive into the research topics of 'Bilingual lexicon extraction using locally weighted linear regression from comparable corpora'. Together they form a unique fingerprint.

Cite this