Skip to main navigation Skip to search Skip to main content

Measuring domain similarity for statistical machine translation

  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

It is well known that the statistical machine translation (SMT) performance suffers when a model is applied to out-of-domain data. It is also known that the more similar the test domain and the training domain are, the more efficient the training data are for SMT performance. Hence, measuring the similarity of domains is an important task to select appropriate training data. The most widely used method uses the cosine similarity function and word frequency. The lack of exploring other approaches motivates us to propose and compare several similarity measures. Aiming for better SMT performance, we compared 10 similarity measures, which are a combination of 2 feature representations and 5 similarity functions. The results show that using the relative word frequency as the feature representation and using the skew divergence as the similarity function performs the best amongst the 10 measures and outperforms random data selection.

Original languageEnglish
Title of host publicationProceedings - 2013 10th International Conference on Fuzzy Systems and Knowledge Discovery, FSKD 2013
PublisherIEEE Computer Society
Pages611-615
Number of pages5
ISBN (Print)9781467352536
DOIs
StatePublished - 2013
Event2013 10th International Conference on Fuzzy Systems and Knowledge Discovery, FSKD 2013 - Shenyang, China
Duration: 23 Jul 201325 Jul 2013

Publication series

NameProceedings - 2013 10th International Conference on Fuzzy Systems and Knowledge Discovery, FSKD 2013

Conference

Conference2013 10th International Conference on Fuzzy Systems and Knowledge Discovery, FSKD 2013
Country/TerritoryChina
CityShenyang
Period23/07/1325/07/13

Keywords

  • domain adaptation
  • domain similarity
  • statistical machine translation(SMT)

Fingerprint

Dive into the research topics of 'Measuring domain similarity for statistical machine translation'. Together they form a unique fingerprint.

Cite this