TY - GEN
T1 - CASIA at SemEval-2022 Task 11
T2 - 16th International Workshop on Semantic Evaluation, SemEval 2022, co-located (hybrid) with The 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics, NAACL 2022
AU - Fu, Jia
AU - Gan, Zhen
AU - Li, Zhucong
AU - Li, Sirui
AU - Sui, Dianbo
AU - Chen, Yubo
AU - Liu, Kang
AU - Zhao, Jun
N1 - Publisher Copyright:
© 2022 Association for Computational Linguistics.
PY - 2022
Y1 - 2022
N2 - This paper describes our approach to develop a complex named entity recognition system in SemEval 2022 Task 11: MultiCoNER Multilingual Complex Named Entity Recognition,Track 9 - Chinese. In this task, we need to identify the entity boundaries and category labels for the six identified categories of CW, LOC, PER, GRP, CORP, and PORD.The task focuses on detecting semantically ambiguous and complex entities in short and low-context settings. We constructed a hybrid system based on Roberta-large model with three training mechanisms and a series of data augmentation. Three training mechanisms include adversarial training, Child-Tuning training, and continued pre-training. The core idea of the hybrid system is to improve the performance of the model in complex environments by introducing more domain knowledge through data augmentation and continuing pre-training domain adaptation of the model. Our proposed method in this paper achieves a macro-F1 of 0.797 on the final test set, ranking second.
AB - This paper describes our approach to develop a complex named entity recognition system in SemEval 2022 Task 11: MultiCoNER Multilingual Complex Named Entity Recognition,Track 9 - Chinese. In this task, we need to identify the entity boundaries and category labels for the six identified categories of CW, LOC, PER, GRP, CORP, and PORD.The task focuses on detecting semantically ambiguous and complex entities in short and low-context settings. We constructed a hybrid system based on Roberta-large model with three training mechanisms and a series of data augmentation. Three training mechanisms include adversarial training, Child-Tuning training, and continued pre-training. The core idea of the hybrid system is to improve the performance of the model in complex environments by introducing more domain knowledge through data augmentation and continuing pre-training domain adaptation of the model. Our proposed method in this paper achieves a macro-F1 of 0.797 on the final test set, ranking second.
UR - https://www.scopus.com/pages/publications/85137606543
U2 - 10.18653/v1/2022.semeval-1.208
DO - 10.18653/v1/2022.semeval-1.208
M3 - 会议稿件
AN - SCOPUS:85137606543
T3 - SemEval 2022 - 16th International Workshop on Semantic Evaluation, Proceedings of the Workshop
SP - 1518
EP - 1523
BT - SemEval 2022 - 16th International Workshop on Semantic Evaluation, Proceedings of the Workshop
A2 - Emerson, Guy
A2 - Schluter, Natalie
A2 - Stanovsky, Gabriel
A2 - Kumar, Ritesh
A2 - Palmer, Alexis
A2 - Schneider, Nathan
A2 - Singh, Siddharth
A2 - Ratan, Shyam
PB - Association for Computational Linguistics (ACL)
Y2 - 14 July 2022 through 15 July 2022
ER -