TY - GEN
T1 - Resolving Chinese zero pronoun with word embedding
AU - Liu, Bingquan
AU - Du, Xinkai
AU - Liu, Ming
AU - Sun, Chengjie
AU - Zheng, Guidong
AU - Zou, Chao
N1 - Publisher Copyright:
© 2018, Springer International Publishing AG.
PY - 2018
Y1 - 2018
N2 - Elliptical sentences are frequently seen in Chinese, especially in some particular situations, such as dialogues, which is challengeable to understand specific semantic. Chinese zero pronoun resolution, which recovers a noun phrase in the elliptical position, is an effective method to help machines understand natural languages. Traditional methods use the features, which are extracted from syntactic parsing trees manually. However, the long running time and the inaccuracy of automatic parsing algorithms have a bad influence on practical applications. In this work, we propose a new method based on long-short-term memory network that calculates dense vector representations for mention pairs without using features from syntactic parsing trees. These representations, which capture significant semantics for zero pronoun resolution, are built on distributed representation of words in surrounding contexts and candidate antecedents. Our method contributes to reducing the manual work of extracting features from parsing tress, which improves the F1-score of Chinese zero pronoun resolution system. Experimental results on OnotoNotes 5.0 Chinese dataset show our method achieves better performance compared with the state-of-the-art method.
AB - Elliptical sentences are frequently seen in Chinese, especially in some particular situations, such as dialogues, which is challengeable to understand specific semantic. Chinese zero pronoun resolution, which recovers a noun phrase in the elliptical position, is an effective method to help machines understand natural languages. Traditional methods use the features, which are extracted from syntactic parsing trees manually. However, the long running time and the inaccuracy of automatic parsing algorithms have a bad influence on practical applications. In this work, we propose a new method based on long-short-term memory network that calculates dense vector representations for mention pairs without using features from syntactic parsing trees. These representations, which capture significant semantics for zero pronoun resolution, are built on distributed representation of words in surrounding contexts and candidate antecedents. Our method contributes to reducing the manual work of extracting features from parsing tress, which improves the F1-score of Chinese zero pronoun resolution system. Experimental results on OnotoNotes 5.0 Chinese dataset show our method achieves better performance compared with the state-of-the-art method.
KW - Chinese zero pronoun resolution
KW - Deep learning
KW - Distributed representation
KW - Long-short-term memory network
UR - https://www.scopus.com/pages/publications/85041188100
U2 - 10.1007/978-3-319-73618-1_72
DO - 10.1007/978-3-319-73618-1_72
M3 - 会议稿件
AN - SCOPUS:85041188100
SN - 9783319736174
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 828
EP - 838
BT - Natural Language Processing and Chinese Computing - 6th CCF International Conference, NLPCC 2017, Proceedings
A2 - Huang, Xuanjing
A2 - Jiang, Jing
A2 - Zhao, Dongyan
A2 - Feng, Yansong
A2 - Hong, Yu
PB - Springer Verlag
T2 - 6th CCF International Conference on Natural Language Processing and Chinese Computing, NLPCC 2017
Y2 - 8 November 2017 through 12 November 2017
ER -