TY - GEN
T1 - HITSZ’s End-To-End Speech Translation Systems Combining Sequence-to-Sequence Auto Speech Recognition Model and Indic Large Language Model for IWSLT 2025 in Indic Track
AU - Wei, Xuchen
AU - Wu, Yangxin
AU - Zhang, Yaoyin
AU - Liu, Henglyu
AU - Chen, Kehai
AU - Bai, Xuefeng
AU - Zhang, Min
N1 - Publisher Copyright:
© 2025 Association for Computational Linguistics
PY - 2025
Y1 - 2025
N2 - This paper presents HITSZ’s submission for the IWSLT 2025 Indic track, focusing on speech-to-text translation (ST) for English-to-Indic and Indic-to-English language pairs. To enhance translation quality in this low-resource scenario, we propose an end-to-end system integrating the pre-trained Whisper automated speech recognition (ASR) model with Krutrim, an Indic-specialized large language model (LLM). Experimental results demonstrate that our end-to-end system achieved average BLEU scores of 28.88 for English-to-Indic directions and 27.86 for Indic-to-English directions. Furthermore, we investigated the Chain-of-Thought (CoT) method. While this method showed potential for significant translation quality improvements on successfully parsed outputs (e.g. a 13.84 BLEU increase for Tamil-to-English), we observed challenges in ensuring the model consistently adheres to the required CoT output format.
AB - This paper presents HITSZ’s submission for the IWSLT 2025 Indic track, focusing on speech-to-text translation (ST) for English-to-Indic and Indic-to-English language pairs. To enhance translation quality in this low-resource scenario, we propose an end-to-end system integrating the pre-trained Whisper automated speech recognition (ASR) model with Krutrim, an Indic-specialized large language model (LLM). Experimental results demonstrate that our end-to-end system achieved average BLEU scores of 28.88 for English-to-Indic directions and 27.86 for Indic-to-English directions. Furthermore, we investigated the Chain-of-Thought (CoT) method. While this method showed potential for significant translation quality improvements on successfully parsed outputs (e.g. a 13.84 BLEU increase for Tamil-to-English), we observed challenges in ensuring the model consistently adheres to the required CoT output format.
UR - https://www.scopus.com/pages/publications/105040083357
U2 - 10.18653/v1/2025.iwslt-1.43
DO - 10.18653/v1/2025.iwslt-1.43
M3 - 会议稿件
AN - SCOPUS:105040083357
T3 - IWSLT 2025 - 22nd International Conference on Spoken Language Translation, Proceedings of the Conference
SP - 405
EP - 411
BT - IWSLT 2025 - 22nd International Conference on Spoken Language Translation, Proceedings of the Conference
A2 - Salesky, Elizabeth
A2 - Federico, Marcello
A2 - Anastasopoulos, Antonis
PB - Association for Computational Linguistics (ACL)
T2 - 22nd International Conference on Spoken Language Translation, IWSLT 2025
Y2 - 31 July 2025 through 1 August 2025
ER -