Skip to main navigation Skip to search Skip to main content

Generating Chinese Named Entity Data from a Parallel Corpus

  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Annotating Named Entity Recognition (NER) training corpora is a costly process but necessary for supervised NER systems. This paper presents an approach to generate large-scale Chinese NER training data from an English-Chinese discourse level aligned parallel corpus. Difficulty of NER is different among languages due to their unique features. For example, the performance of English NER systems is usually higher than the Chinese ones on average. In our method, we first employ a high performance NER system on one side of a bilingual corpus. And then, we project the NE labels to the other side according to the word level alignment. At last, we select high-quality labeled sentences using different strategies and generate an NER training corpus. In our experiments, we generate a Chinese NER corpus with 167,100 sentences through an English-Chinese parallel corpus. The system trained on the automatically generated corpus attains a comparable result with the one trained on the manually-annotated corpus. Further experiments show that the NER performance is significantly improved on two different evaluation sets by using the generated training data as an additional corpus to the manually-labeled data.

Original languageEnglish
Title of host publicationIJCNLP 2011 - Proceedings of the 5th International Joint Conference on Natural Language Processing
EditorsHaifeng Wang, David Yarowsky
PublisherAssociation for Computational Linguistics (ACL)
Pages264-272
Number of pages9
ISBN (Electronic)9789744665645
StatePublished - 2011
Externally publishedYes
Event5th International Joint Conference on Natural Language Processing, IJCNLP 2011 - Chiang Mai, Thailand
Duration: 8 Nov 201113 Nov 2011

Publication series

NameIJCNLP 2011 - Proceedings of the 5th International Joint Conference on Natural Language Processing

Conference

Conference5th International Joint Conference on Natural Language Processing, IJCNLP 2011
Country/TerritoryThailand
CityChiang Mai
Period8/11/1113/11/11

Fingerprint

Dive into the research topics of 'Generating Chinese Named Entity Data from a Parallel Corpus'. Together they form a unique fingerprint.

Cite this