Skip to main navigation Skip to search Skip to main content

Exploiting Translation Model for Parallel Corpus Mining

  • Chongman Leong*
  • , Xuebo Liu
  • , Derek F. Wong
  • , Lidia S. Chao
  • *Corresponding author for this work
  • University of Macau

Research output: Contribution to journalArticlepeer-review

Abstract

Parallel corpus mining (PCM) is beneficial for many corpus-based natural language processing tasks, e.g., machine translation and bilingual dictionary induction, especially in low-resource languages and domains. It relies heavily on cross-lingual representations to model the interdependencies between different languages and determine whether sentences are parallel or not. In this paper, we take the first step towards exploiting the multilingual Transformer translation model to produce expressive sentence representations for PCM. Since the traditional Transformer lacks an immediate sentence representation, we pool the output representation of the encoder as the sentence representation, which is further optimized as a part of the training flow of the translation model. Experiments conducted on the BUCC PCM task show that the proposed method improves mining performance over the existing methods with the assistance of the pre-trained multilingual BERT. To further test the usability of the proposed method, we mine parallel sentences from public resources and find that the mined sentences can indeed enhance low-resource machine translation.

Original languageEnglish
Article number9516978
Pages (from-to)2829-2839
Number of pages11
JournalIEEE/ACM Transactions on Audio Speech and Language Processing
Volume29
DOIs
StatePublished - 2021
Externally publishedYes

Keywords

  • Chinese-Portuguese translation
  • neural machine translation
  • parallel corpus mining

Fingerprint

Dive into the research topics of 'Exploiting Translation Model for Parallel Corpus Mining'. Together they form a unique fingerprint.

Cite this