Skip to main navigation Skip to search Skip to main content

Extracting parallel phrases from comparable corpora

  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The state-of-the-art statistical machine translation models are trained with the parallel corpora. However, the traditional SMT loses its power when it comes to language pairs with few bilingual resources. This paper proposes a novel method that treats the phrase extraction as a classification task. We first automatically generate the training and testing phrase pairs for the classifier. Then, we train a SVM classifier which can determine the phrase pairs are either parallel or non-parallel. The proposed approach is evaluated on the translation task of Chinese-English. Experimental results show that the precision of the classifier on test sets is above 70% and the accuracy is above 98% The quality of the extracted data is also evaluated by measuring the impact on the performance of a state-of-the-art SMT system, which is built with a small parallel corpus. It shows better results over the baseline system.

Original languageEnglish
Title of host publicationProceedings of the International Conference on Asian Language Processing 2014, IALP 2014
EditorsRafael E. Banchs, Minghui Dong, Yanfeng Lu, Bali Ranaivo-Malancon
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages166-169
Number of pages4
ISBN (Electronic)9781479953301
DOIs
StatePublished - 3 Dec 2014
Externally publishedYes
EventInternational Conference on Asian Language Processing 2014, IALP 2014 - Kuching, Malaysia
Duration: 20 Oct 201422 Oct 2014

Publication series

NameProceedings of the International Conference on Asian Language Processing 2014, IALP 2014

Conference

ConferenceInternational Conference on Asian Language Processing 2014, IALP 2014
Country/TerritoryMalaysia
CityKuching
Period20/10/1422/10/14

Keywords

  • Statistical Machine Translation
  • Support Vector Machine
  • classification
  • comparable corpus

Fingerprint

Dive into the research topics of 'Extracting parallel phrases from comparable corpora'. Together they form a unique fingerprint.

Cite this