TY - GEN
T1 - A mixed model for cross lingual opinion analysis
AU - Gui, Lin
AU - Xu, Ruifeng
AU - Xu, Jun
AU - Yuan, Li
AU - Yao, Yuanlin
AU - Zhou, Jiyun
AU - Qiu, Qiaoyun
AU - Wang, Shuwei
AU - Wong, Kam Fai
AU - Cheung, Ricky
PY - 2013
Y1 - 2013
N2 - The performances of machine learning based opinion analysis systems are always puzzled by the insufficient training opinion corpus. Such problem becomes more serious for the resource-poor languages. Thus, the cross-lingual opinion analysis (CLOA) technique, which leverages opinion resources on one (source) language to another (target) language for improving the opinion analysis on target language, attracts more research interests. Currently, the transfer learning based CLOA approach sometimes falls to over fitting on single language resource, while the performance of the co-training based CLOA approach always achieves limited improvement during bi-lingual decision. Target to these problems, in this study, we propose a mixed CLOA model, which estimates the confidence of each monolingual opinion analysis system by using their training errors through bilingual transfer self-training and co-training, respectively. By using the weighted average distances between samples and classification hyper-planes as the confidence, the opinion polarity of testing samples are classified. The evaluations on NLP&CC 2013 CLOA bakeoff dataset show that this approach achieves the best performance, which outperforms transfer learning and co-training based approaches.
AB - The performances of machine learning based opinion analysis systems are always puzzled by the insufficient training opinion corpus. Such problem becomes more serious for the resource-poor languages. Thus, the cross-lingual opinion analysis (CLOA) technique, which leverages opinion resources on one (source) language to another (target) language for improving the opinion analysis on target language, attracts more research interests. Currently, the transfer learning based CLOA approach sometimes falls to over fitting on single language resource, while the performance of the co-training based CLOA approach always achieves limited improvement during bi-lingual decision. Target to these problems, in this study, we propose a mixed CLOA model, which estimates the confidence of each monolingual opinion analysis system by using their training errors through bilingual transfer self-training and co-training, respectively. By using the weighted average distances between samples and classification hyper-planes as the confidence, the opinion polarity of testing samples are classified. The evaluations on NLP&CC 2013 CLOA bakeoff dataset show that this approach achieves the best performance, which outperforms transfer learning and co-training based approaches.
KW - Co-training
KW - Cross lingual opinion analysis
KW - Mixed model
KW - Transfer self-training
UR - https://www.scopus.com/pages/publications/84901488468
U2 - 10.1007/978-3-642-41644-6_10
DO - 10.1007/978-3-642-41644-6_10
M3 - 会议稿件
AN - SCOPUS:84901488468
SN - 9783642416439
T3 - Communications in Computer and Information Science
SP - 93
EP - 104
BT - Natural Language Processing and Chinese Computing - Second CCF Conference, NLPCC 2013, Proceedings
PB - Springer Verlag
T2 - 2nd CCF Conference on Natural Language Processing and Chinese Computing, NLPCC 2013
Y2 - 15 November 2013 through 19 November 2013
ER -