TY - GEN
T1 - Generating Chinese Named Entity Data from a Parallel Corpus
AU - Fu, Ruiji
AU - Qin, Bing
AU - Liu, Ting
N1 - Publisher Copyright:
© 2011 AFNLP
PY - 2011
Y1 - 2011
N2 - Annotating Named Entity Recognition (NER) training corpora is a costly process but necessary for supervised NER systems. This paper presents an approach to generate large-scale Chinese NER training data from an English-Chinese discourse level aligned parallel corpus. Difficulty of NER is different among languages due to their unique features. For example, the performance of English NER systems is usually higher than the Chinese ones on average. In our method, we first employ a high performance NER system on one side of a bilingual corpus. And then, we project the NE labels to the other side according to the word level alignment. At last, we select high-quality labeled sentences using different strategies and generate an NER training corpus. In our experiments, we generate a Chinese NER corpus with 167,100 sentences through an English-Chinese parallel corpus. The system trained on the automatically generated corpus attains a comparable result with the one trained on the manually-annotated corpus. Further experiments show that the NER performance is significantly improved on two different evaluation sets by using the generated training data as an additional corpus to the manually-labeled data.
AB - Annotating Named Entity Recognition (NER) training corpora is a costly process but necessary for supervised NER systems. This paper presents an approach to generate large-scale Chinese NER training data from an English-Chinese discourse level aligned parallel corpus. Difficulty of NER is different among languages due to their unique features. For example, the performance of English NER systems is usually higher than the Chinese ones on average. In our method, we first employ a high performance NER system on one side of a bilingual corpus. And then, we project the NE labels to the other side according to the word level alignment. At last, we select high-quality labeled sentences using different strategies and generate an NER training corpus. In our experiments, we generate a Chinese NER corpus with 167,100 sentences through an English-Chinese parallel corpus. The system trained on the automatically generated corpus attains a comparable result with the one trained on the manually-annotated corpus. Further experiments show that the NER performance is significantly improved on two different evaluation sets by using the generated training data as an additional corpus to the manually-labeled data.
UR - https://www.scopus.com/pages/publications/84870292816
M3 - 会议稿件
AN - SCOPUS:84870292816
T3 - IJCNLP 2011 - Proceedings of the 5th International Joint Conference on Natural Language Processing
SP - 264
EP - 272
BT - IJCNLP 2011 - Proceedings of the 5th International Joint Conference on Natural Language Processing
A2 - Wang, Haifeng
A2 - Yarowsky, David
PB - Association for Computational Linguistics (ACL)
T2 - 5th International Joint Conference on Natural Language Processing, IJCNLP 2011
Y2 - 8 November 2011 through 13 November 2011
ER -