TY - GEN
T1 - CCL23-Eval任务7赛道一系统报告:Suda &Alibaba 文本纠错系统
AU - Jiang, Haochen
AU - Liu, Yumeng
AU - Zhou, Houquan
AU - Qiao, Ziheng
AU - Zhang, Bo
AU - Li, Chen
AU - Li, Zhenghua
AU - Zhang, Min
N1 - Publisher Copyright:
© 2023 China National Conference on Computational Linguistics.
PY - 2023
Y1 - 2023
N2 - The article describes the submission of Suda &Alibaba Error Correction Team for Track 1 of the Multidimensional Chinese Learner Text Correction (CCL2023) evaluation task. In terms of models, we used both sequence-to-sequence and sequence-to-edit correction models. For data, we conducted a three-stage training using pseudo data constructed based on confusion sets, real data from Lang-8, and the development set from YACLC. In the open task, we also utilized additional data such as HSK and CGED for training. We employed a series of effective performance enhancement techniques, including rule-based data augmentation, data cleaning, post-processing, and model ensembling. Moreover, we explored the use of large models such as GPT3.5 and GPT4 to assist Chinese text correction and tried various prompts. In both the closed and open tasks, our team ranked first in minimum edits, fluency improvement, and average F0.5 scores.
AB - The article describes the submission of Suda &Alibaba Error Correction Team for Track 1 of the Multidimensional Chinese Learner Text Correction (CCL2023) evaluation task. In terms of models, we used both sequence-to-sequence and sequence-to-edit correction models. For data, we conducted a three-stage training using pseudo data constructed based on confusion sets, real data from Lang-8, and the development set from YACLC. In the open task, we also utilized additional data such as HSK and CGED for training. We employed a series of effective performance enhancement techniques, including rule-based data augmentation, data cleaning, post-processing, and model ensembling. Moreover, we explored the use of large models such as GPT3.5 and GPT4 to assist Chinese text correction and tried various prompts. In both the closed and open tasks, our team ranked first in minimum edits, fluency improvement, and average F0.5 scores.
KW - Sequence-to-edit
KW - Sequence-to-sequence
KW - Text Correction
UR - https://www.scopus.com/pages/publications/85175979916
M3 - 会议稿件
AN - SCOPUS:85175979916
T3 - Proceedings of the 22nd Chinese National Conference on Computational Linguistics, CCL 2023
SP - 220
EP - 229
BT - Evaluations
A2 - Sun, Maosong
A2 - Qin, Bing
A2 - Qiu, Xipeng
A2 - Jiang, Jing
A2 - Han, Xianpei
PB - Association for Computational Linguistics (ACL)
T2 - 22nd Chinese National Conference on Computational Linguistics, CCL 2023
Y2 - 3 August 2023 through 5 August 2023
ER -