TY - GEN
T1 - Context Consistency between Training and Inference in Simultaneous Machine Translation
AU - Zhong, Meizhi
AU - Liu, Lemao
AU - Chen, Kehai
AU - Yang, Mingming
AU - Zhang, Min
N1 - Publisher Copyright:
© 2024 Association for Computational Linguistics.
PY - 2024
Y1 - 2024
N2 - Simultaneous Machine Translation (SiMT) aims to yield a real-time partial translation with a monotonically growing source-side context. However, there is a counterintuitive phenomenon about the context usage between training and inference: e.g., in wait-k inference, model consistently trained with wait-k is much worse than that model inconsistently trained with wait-k' (k' ? k) in terms of translation quality. To this end, we first investigate the underlying reasons behind this phenomenon and uncover the following two factors: 1) the limited correlation between translation quality and training loss; 2) exposure bias between training and inference. Based on both reasons, we then propose an effective training approach called context consistency training accordingly, which encourages consistent context usage between training and inference by optimizing translation quality and latency as bi-objectives and exposing the predictions to the model during the training. The experiments on three language pairs demonstrate that our SiMT system encouraging context consistency outperforms existing SiMT systems with context inconsistency for the first time.
AB - Simultaneous Machine Translation (SiMT) aims to yield a real-time partial translation with a monotonically growing source-side context. However, there is a counterintuitive phenomenon about the context usage between training and inference: e.g., in wait-k inference, model consistently trained with wait-k is much worse than that model inconsistently trained with wait-k' (k' ? k) in terms of translation quality. To this end, we first investigate the underlying reasons behind this phenomenon and uncover the following two factors: 1) the limited correlation between translation quality and training loss; 2) exposure bias between training and inference. Based on both reasons, we then propose an effective training approach called context consistency training accordingly, which encourages consistent context usage between training and inference by optimizing translation quality and latency as bi-objectives and exposing the predictions to the model during the training. The experiments on three language pairs demonstrate that our SiMT system encouraging context consistency outperforms existing SiMT systems with context inconsistency for the first time.
UR - https://www.scopus.com/pages/publications/85204472463
U2 - 10.18653/v1/2024.acl-long.727
DO - 10.18653/v1/2024.acl-long.727
M3 - 会议稿件
AN - SCOPUS:85204472463
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 13465
EP - 13476
BT - Long Papers
A2 - Ku, Lun-Wei
A2 - Martins, Andre F. T.
A2 - Srikumar, Vivek
PB - Association for Computational Linguistics (ACL)
T2 - 62nd Annual Meeting of the Association for Computational Linguistics, ACL 2024
Y2 - 11 August 2024 through 16 August 2024
ER -