Abstract
The annotation error detection (AED) task aims to automatically identify annotation errors in a dataset, which is crucial for ensuring the reliability and effectiveness of expert and intelligent systems across diverse applications. Most previous works either employ synthesized data, or subset of crowdsourced datasets. In contrast, this work focuses on detecting errors in painstakingly annotated data, using part-of-speech (POS) tagging as a case study. We construct a high-quality Chinese AED dataset, named CTB7E, by manually re-annotating the test set of CTB7. Among 81,578 tags, we identify approximately 1,700 erroneous tags, resulting in a 2.1 % error rate. We for the first time apply Kullback-Leibler (KL) divergence to AED and propose two new metrics. We investigate a wide range of AED approaches on both CTB7E and a synthesized dataset, under both single-model and Monte Carlo dropout settings. The results and analyses reveal interesting insights. We will release our data and code at https://github.com/yahui19960717/POS_AED.git to facilitate further research and collaboration in this area.
| Original language | English |
|---|---|
| Article number | 128374 |
| Journal | Expert Systems with Applications |
| Volume | 290 |
| DOIs | |
| State | Published - 25 Sep 2025 |
| Externally published | Yes |
Keywords
- Annotation error detection
- Chinese part-of-speech
- Dataset
- KL divergence
Fingerprint
Dive into the research topics of 'Annotation error detection in painstakingly annotated data: Part-of-speech tagging as a case study'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver