Skip to main navigation Skip to search Skip to main content

Annotation error detection in painstakingly annotated data: Part-of-speech tagging as a case study

  • Yahui Liu
  • , Zhenghua Li
  • , Chen Gong*
  • , Shilin Zhou
  • , Min Zhang
  • *Corresponding author for this work
  • Soochow University

Research output: Contribution to journalArticlepeer-review

Abstract

The annotation error detection (AED) task aims to automatically identify annotation errors in a dataset, which is crucial for ensuring the reliability and effectiveness of expert and intelligent systems across diverse applications. Most previous works either employ synthesized data, or subset of crowdsourced datasets. In contrast, this work focuses on detecting errors in painstakingly annotated data, using part-of-speech (POS) tagging as a case study. We construct a high-quality Chinese AED dataset, named CTB7E, by manually re-annotating the test set of CTB7. Among 81,578 tags, we identify approximately 1,700 erroneous tags, resulting in a 2.1 % error rate. We for the first time apply Kullback-Leibler (KL) divergence to AED and propose two new metrics. We investigate a wide range of AED approaches on both CTB7E and a synthesized dataset, under both single-model and Monte Carlo dropout settings. The results and analyses reveal interesting insights. We will release our data and code at https://github.com/yahui19960717/POS_AED.git to facilitate further research and collaboration in this area.

Original languageEnglish
Article number128374
JournalExpert Systems with Applications
Volume290
DOIs
StatePublished - 25 Sep 2025
Externally publishedYes

Keywords

  • Annotation error detection
  • Chinese part-of-speech
  • Dataset
  • KL divergence

Fingerprint

Dive into the research topics of 'Annotation error detection in painstakingly annotated data: Part-of-speech tagging as a case study'. Together they form a unique fingerprint.

Cite this