Skip to main navigation Skip to search Skip to main content

Combining sentence length with location information to align monolingual parallel texts

  • Weigang Li*
  • , Ting Liu
  • , Sheng Li
  • *Corresponding author for this work
  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Contribution to journalConference articlepeer-review

Abstract

Abundant Chinese paraphrasing resource on Internet can be attained from different Chinese translations of one foreign masterpiece. Paraphrases corpus is the corpus that includes sentence pairs to convey the same information. The irregular characteristics of the real monolingual parallel texts, especially without the strictly aligned paragraph boundaries between two translations, bring a challenge to alignment technology. The traditional alignment methods on bilingual texts have some difficulties in competency for doing this. A new method for aligning real monolingual parallel texts using sentence pair's length and location information is described in this paper. The model was motivated by the observation that the location of a sentence pair with certain length is distributed in the whole text similarly. And presently, a paraphrases corpus with about fifty thousand sentence pairs is constructed.

Original languageEnglish
Pages (from-to)118-128
Number of pages11
JournalLecture Notes in Computer Science
Volume3411
DOIs
StatePublished - 2005
Externally publishedYes
EventAsia Information Retrieval Symposium, AIRS 2004 - Beijing, China
Duration: 18 Oct 200420 Oct 2004

Fingerprint

Dive into the research topics of 'Combining sentence length with location information to align monolingual parallel texts'. Together they form a unique fingerprint.

Cite this