Skip to main navigation Skip to search Skip to main content

Grammar comparison study for Translational Equivalence Modeling and Statistical Machine Translation

  • Min Zhang*
  • , Hongfei Jiang
  • , Haizhou Li
  • , Aiti Aw
  • , Sheng Li
  • *Corresponding author for this work
  • Agency for Science, Technology and Research, Singapore
  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

This paper presents a general platform, namely synchronous tree sequence substitution grammar (STSSG), for the grammar comparison study in Translational Equivalence Modeling (TEM) and Statistical Machine Translation (SMT). Under the STSSG platform, we compare the expressive abilities of various grammars through synchronous parsing and a real translation platform on a variety of Chinese-English bilingual corpora. Experimental results show that the STSSG is able to better explain the data in parallel corpora than other grammars. Our study further finds that the complexity of structure divergence is much higher than suggested in literature, which imposes a big challenge to syntactic transformation-based SMT.

Original languageEnglish
Title of host publicationColing 2008 - 22nd International Conference on Computational Linguistics, Proceedings of the Conference
PublisherAssociation for Computational Linguistics (ACL)
Pages1097-1104
Number of pages8
ISBN (Print)9781905593446
DOIs
StatePublished - 2008
Externally publishedYes
Event22nd International Conference on Computational Linguistics, Coling 2008 - Manchester, United Kingdom
Duration: 18 Aug 200822 Aug 2008

Publication series

NameColing 2008 - 22nd International Conference on Computational Linguistics, Proceedings of the Conference
Volume1

Conference

Conference22nd International Conference on Computational Linguistics, Coling 2008
Country/TerritoryUnited Kingdom
CityManchester
Period18/08/0822/08/08

Fingerprint

Dive into the research topics of 'Grammar comparison study for Translational Equivalence Modeling and Statistical Machine Translation'. Together they form a unique fingerprint.

Cite this