Challenge: Existing approaches to speech-to-speech translation rely on cascaded pipelines . current approaches rely only on text representations, but they suffer from errors and latency . a new direct speech translation framework is proposed to bridge linguistic gaps .
Approach: They propose a sequence-to-sequence direct speech translation framework that can translate speech from one Indian language to another without relying on intermediate text representations.
Outcome: The proposed framework can translate speech from one Indian language to another without relying on intermediate text representations.

Similar Papers

Tutorial: End-to-End Speech Translation (2021.eacl-tutorials)

Copied to clipboard

Challenge: Speech translation is the translation of speech in one language typically to text in another, traditionally accomplished through a combination of automatic speech recognition and machine translation.
Approach: This tutorial introduces the techniques used in cutting-edge research on speech translation.
Outcome: The proposed models achieve state-of-the-art performance with end-to-end speech translation for both high- and low-resource languages.
Speech Translation and the End-to-End Promise: Taking Stock of Where We Are (2020.acl-main)

Copied to clipboard

Challenge: Until recently, the only feasible approach to translating acoustic speech signals into text was the cascaded approach.
Approach: They propose a classification of the main challenges of traditional approaches to speech translation . they argue that end-to-end models fall short due to compromises made to address data scarcity .
Outcome: This paper provides a brief survey of the main challenges of traditional approaches in speech translation . it reveals that many end-to-end models fail due to compromises made to address data scarcity.
Towards Speech to Speech Machine Translation focusing on Indian Languages (2023.eacl-demo)

Copied to clipboard

Challenge: SSMT is a web application for translating videos from one language to another by cascading multiple language modules.
Approach: They introduce an SSMT pipeline for translating videos from one language to another by cascading multiple language modules.
Outcome: The proposed system can get 3.5+ MOS score for English to Hindi using human intervention.
LLaST: Improved End-to-end Speech Translation System Leveraged by Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: ***LLaST*** is a framework for building high-performance Large Language model based Speech-to-text Translation systems.
Approach: They propose a framework for building high-performance Large Language model based Speech-to-text Translation systems.
Outcome: The proposed model outperforms the CoVoST-2 benchmark and showcases exceptional scaling capabilities powered by LLMs.
Phone Features Improve Speech Translation (2020.acl-main)

Copied to clipboard

Challenge: End-to-end models for speech translation more tightly couple speech recognition (ASR) and machine translation (MT) compared to cascades, but performance gap remains in low-resource conditions .
Approach: They propose two methods to incorporate phone features into current neural speech translation models.
Outcome: The proposed models outperform existing models and cascades by up to 9 BLEU on low-resource conditions.
CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving (2025.coling-main)

Copied to clipboard

Challenge: More than half of the world's population is presumed to be bilingual . spoken translation of code-switched speech has been under-explored .
Approach: They propose an end-to-end model architecture CoSTA that scaffolds on pretrained ASR and MT modules.
Outcome: The proposed model outperforms existing models by 3.5 BLEU points in spoken translation of code-switched speech.
Streaming Models for Joint Speech Recognition and Translation (2021.eacl-main)

Copied to clipboard

Challenge: Using end-to-end models for speech translation has become a focus of the ST community . cascaded models have the advantage of including automatic speech recognition output .
Approach: They propose a model that condenses sound waves into translated text and integrates automatic speech recognition outputs into the models.
Outcome: The proposed model is statistically similar to cascading models, but has half the number of parameters.
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation (2023.emnlp-demo)

Copied to clipboard

Challenge: a framework to evaluate low-latency speech translations is currently only limited to specific aspects and is not able to compare different approaches.
Approach: They propose a framework to perform and evaluate low-latency speech translation in realistic conditions.
Outcome: The proposed framework evaluates various aspects of low-latency speech translation under realistic conditions.
Towards Building Large Scale Datasets and State-of-the-Art Automatic Speech Translation Systems for 14 Indian Languages (2025.acl-long)

Copied to clipboard

Challenge: Existing datasets that cover only a fraction of Indian languages lack the breadth needed to generalize beyond curated benchmarks.
Approach: They propose to build the largest speech translation dataset for Indian languages . they use a three-step methodology to gather data and train a model that performs better .
Outcome: The proposed model improves on existing models and is open-source with permissive licenses.
Large-scale Machine Translation for Indian Languages in E-commerce under Low Resource Constraints (2022.emnlp-industry)

Copied to clipboard

Challenge: We have deployed reliable and precise large-scale machine translation systems for several Indian regional languages.
Approach: They develop a structured model development pipeline as a closed feedback loop with external manual feedback through an Active Learning component.
Outcome: The proposed model improves over iterations for English to Hindi and for other languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations