Skip to main navigation Skip to search Skip to main content

Dependency Scoring Learning and Corpus Boosting for Translation-Based Cross-Lingual Dependency Parsing

  • Harbin Institute of Technology Shenzhen
  • Information Center of Ministry of Science and Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Dependency parsing is a fundamental task in natural language processing that involves identifying the grammatical relationships between words in a sentence. One promising approach for performing this task in languages lacking annotated treebanks is treebank translation, which utilizes word alignments to map dependencies from a source treebank to the corresponding target translation. However, due to language differences and the limitations of word alignment tools, this method would inevitably generate noise during mapping. To reduce the effect of noise, we first exploit MetaNet to compute quality scores for each dependency and identify low-score ones as noise. MetaNet is a fake teacher that learns to score homework (dependencies) by comparing answers from the top student (strong parser) and the regular student (weak parser) without knowing the correct answer (gold-standard). With the scoring capability of MetaNet, we design an iterative algorithm to boost the target treebank quality, which trains with high-quality dependencies and relabels the low-quality dependencies. Our method achieves better results than the originally translated treebanks and shows highly competitive performances with prior methods on the Universal Dependency Treebanks v2.2. We also provide detailed analysis and discussions.

Original languageEnglish
Article number83
JournalACM Transactions on Asian and Low-Resource Language Information Processing
Volume24
Issue number8
DOIs
StatePublished - 21 Aug 2025
Externally publishedYes

Keywords

  • Treebank translation
  • cross-lingual
  • dependency parsing
  • noise filtering

Fingerprint

Dive into the research topics of 'Dependency Scoring Learning and Corpus Boosting for Translation-Based Cross-Lingual Dependency Parsing'. Together they form a unique fingerprint.

Cite this