TY - GEN
T1 - Can Diffusion Model Achieve Better Performance in Text Generation? Bridging the Gap between Training and Inference!
AU - Tang, Zecheng
AU - Wang, Pinzheng
AU - Zhou, Keyan
AU - Li, Juntao
AU - Cao, Ziqiang
AU - Zhang, Min
N1 - Publisher Copyright:
© 2023 Association for Computational Linguistics.
PY - 2023
Y1 - 2023
N2 - Diffusion models have been successfully adapted to text generation tasks by mapping the discrete text into the continuous space. However, there exist nonnegligible gaps between training and inference, owing to the absence of the forward process during inference. Thus, the model only predicts based on the previously generated reverse noise rather than the noise computed by the forward process. Besides, the widely-used downsampling strategy in speeding up the inference will cause the mismatch of diffusion trajectories between training and inference. To understand and mitigate the above two types of training-inference discrepancies, we launch a thorough preliminary study. Based on our observations, we propose two simple yet effective methods to bridge the gaps mentioned above, named Distance Penalty and Adaptive Decay Sampling. Extensive experiments on 6 generation tasks confirm the superiority of our methods, which can achieve 100× → 200× speedup with better performance. Our code is available at https://github.com/CODINNLG/Bridge_Gap_Diffusion.
AB - Diffusion models have been successfully adapted to text generation tasks by mapping the discrete text into the continuous space. However, there exist nonnegligible gaps between training and inference, owing to the absence of the forward process during inference. Thus, the model only predicts based on the previously generated reverse noise rather than the noise computed by the forward process. Besides, the widely-used downsampling strategy in speeding up the inference will cause the mismatch of diffusion trajectories between training and inference. To understand and mitigate the above two types of training-inference discrepancies, we launch a thorough preliminary study. Based on our observations, we propose two simple yet effective methods to bridge the gaps mentioned above, named Distance Penalty and Adaptive Decay Sampling. Extensive experiments on 6 generation tasks confirm the superiority of our methods, which can achieve 100× → 200× speedup with better performance. Our code is available at https://github.com/CODINNLG/Bridge_Gap_Diffusion.
UR - https://www.scopus.com/pages/publications/85175445471
U2 - 10.18653/v1/2023.findings-acl.721
DO - 10.18653/v1/2023.findings-acl.721
M3 - 会议稿件
AN - SCOPUS:85175445471
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 11359
EP - 11386
BT - Findings of the Association for Computational Linguistics, ACL 2023
PB - Association for Computational Linguistics (ACL)
T2 - Findings of the Association for Computational Linguistics, ACL 2023
Y2 - 9 July 2023 through 14 July 2023
ER -