TY - GEN
T1 - Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
AU - Du, Yexing
AU - Pan, Youcheng
AU - Ma, Ziyang
AU - Yang, Bo
AU - Yang, Yifan
AU - Deng, Keqi
AU - Chen, Xie
AU - Xiang, Yang
AU - Liu, Ming
AU - Qin, Bing
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics.
PY - 2025
Y1 - 2025
N2 - Multimodal Large Language Models (MLLMs) have achieved significant success in Speech-to-Text Translation (S2TT) tasks. While most existing research has focused on English-centric translation directions, the exploration of many-to-many translation is still limited by the scarcity of parallel data. To address this, we propose a three-stage curriculum learning strategy that leverages the machine translation capabilities of large language models and adapts them to S2TT tasks, enabling effective learning in low-resource settings. We trained MLLMs with varying parameter sizes (3B, 7B, and 32B) and evaluated the proposed strategy using the FLEURS and CoVoST-2 datasets. Experimental results show that the proposed strategy achieves state-of-the-art average performance in 15 × 14 language pairs, requiring fewer than 10 hours of speech data per language to achieve competitive results.
AB - Multimodal Large Language Models (MLLMs) have achieved significant success in Speech-to-Text Translation (S2TT) tasks. While most existing research has focused on English-centric translation directions, the exploration of many-to-many translation is still limited by the scarcity of parallel data. To address this, we propose a three-stage curriculum learning strategy that leverages the machine translation capabilities of large language models and adapts them to S2TT tasks, enabling effective learning in low-resource settings. We trained MLLMs with varying parameter sizes (3B, 7B, and 32B) and evaluated the proposed strategy using the FLEURS and CoVoST-2 datasets. Experimental results show that the proposed strategy achieves state-of-the-art average performance in 15 × 14 language pairs, requiring fewer than 10 hours of speech data per language to achieve competitive results.
UR - https://www.scopus.com/pages/publications/105021018318
U2 - 10.18653/v1/2025.acl-long.610
DO - 10.18653/v1/2025.acl-long.610
M3 - 会议稿件
AN - SCOPUS:105021018318
T3 - Proceedings of the Annual Meeting of the Association for Computational Linguistics
SP - 12466
EP - 12478
BT - Long Papers
A2 - Che, Wanxiang
A2 - Nabende, Joyce
A2 - Shutova, Ekaterina
A2 - Pilehvar, Mohammad Taher
PB - Association for Computational Linguistics (ACL)
T2 - 63rd Annual Meeting of the Association for Computational Linguistics, ACL 2025
Y2 - 27 July 2025 through 1 August 2025
ER -