Skip to main navigation Skip to search Skip to main content

HITSZ’s End-To-End Speech Translation Systems Combining Sequence-to-Sequence Auto Speech Recognition Model and Indic Large Language Model for IWSLT 2025 in Indic Track

  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

This paper presents HITSZ’s submission for the IWSLT 2025 Indic track, focusing on speech-to-text translation (ST) for English-to-Indic and Indic-to-English language pairs. To enhance translation quality in this low-resource scenario, we propose an end-to-end system integrating the pre-trained Whisper automated speech recognition (ASR) model with Krutrim, an Indic-specialized large language model (LLM). Experimental results demonstrate that our end-to-end system achieved average BLEU scores of 28.88 for English-to-Indic directions and 27.86 for Indic-to-English directions. Furthermore, we investigated the Chain-of-Thought (CoT) method. While this method showed potential for significant translation quality improvements on successfully parsed outputs (e.g. a 13.84 BLEU increase for Tamil-to-English), we observed challenges in ensuring the model consistently adheres to the required CoT output format.

Original languageEnglish
Title of host publicationIWSLT 2025 - 22nd International Conference on Spoken Language Translation, Proceedings of the Conference
EditorsElizabeth Salesky, Marcello Federico, Antonis Anastasopoulos
PublisherAssociation for Computational Linguistics (ACL)
Pages405-411
Number of pages7
ISBN (Electronic)9798891762725
DOIs
StatePublished - 2025
Externally publishedYes
Event22nd International Conference on Spoken Language Translation, IWSLT 2025 - Hybrid, Vienna, Austria
Duration: 31 Jul 20251 Aug 2025

Publication series

NameIWSLT 2025 - 22nd International Conference on Spoken Language Translation, Proceedings of the Conference

Conference

Conference22nd International Conference on Spoken Language Translation, IWSLT 2025
Country/TerritoryAustria
CityHybrid, Vienna
Period31/07/251/08/25

Fingerprint

Dive into the research topics of 'HITSZ’s End-To-End Speech Translation Systems Combining Sequence-to-Sequence Auto Speech Recognition Model and Indic Large Language Model for IWSLT 2025 in Indic Track'. Together they form a unique fingerprint.

Cite this