Exploring Phoneme-Level Speech Representations for End-to-End Speech Translation (P19-1)
Copied to clipboard
| Challenge: | Previous work on end-to-end translation from speech uses frame-level features as speech representations, which creates longer, sparser sequences than text. |
| Approach: | They propose a method to generate compressed phoneme-like speech representations that generate shorter, higher-level source sequences for translation. |
| Outcome: | The proposed method improves translation performance by 5 BLEU on high and low resource languages and reduces training time by 60%. |
Similar Papers
Speech Translation and the End-to-End Promise: Taking Stock of Where We Are (2020.acl-main)
Copied to clipboard
| Challenge: | Until recently, the only feasible approach to translating acoustic speech signals into text was the cascaded approach. |
| Approach: | They propose a classification of the main challenges of traditional approaches to speech translation . they argue that end-to-end models fall short due to compromises made to address data scarcity . |
| Outcome: | This paper provides a brief survey of the main challenges of traditional approaches in speech translation . it reveals that many end-to-end models fail due to compromises made to address data scarcity. |
Tutorial: End-to-End Speech Translation (2021.eacl-tutorials)
Copied to clipboard
| Challenge: | Speech translation is the translation of speech in one language typically to text in another, traditionally accomplished through a combination of automatic speech recognition and machine translation. |
| Approach: | This tutorial introduces the techniques used in cutting-edge research on speech translation. |
| Outcome: | The proposed models achieve state-of-the-art performance with end-to-end speech translation for both high- and low-resource languages. |
Simple and Effective Unsupervised Speech Translation (2023.acl-long)
Copied to clipboard
Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov, Yun Tang, Wei-Ning Hsu, Michael Auli, Juan Pino
| Challenge: | Existing methods to train speech models without labeled data are limited for most languages. |
| Approach: | They propose a pipeline approach to build speech translation systems without labeled data by leveraging recent advances in unsupervised speech recognition, machine translation and speech synthesis. |
| Outcome: | The proposed approach outperforms the state-of-the-art in unsupervised speech recognition by 3.2 BLEU on the Libri-Trans benchmark and the best supervised end-to-end models from only two years ago by an average of 5.0 BLUE over five X-En directions. |
Phone Features Improve Speech Translation (2020.acl-main)
Copied to clipboard
| Challenge: | End-to-end models for speech translation more tightly couple speech recognition (ASR) and machine translation (MT) compared to cascades, but performance gap remains in low-resource conditions . |
| Approach: | They propose two methods to incorporate phone features into current neural speech translation models. |
| Outcome: | The proposed models outperform existing models and cascades by up to 9 BLEU on low-resource conditions. |
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation (2023.emnlp-demo)
Copied to clipboard
Christian Huber, Tu Anh Dinh, Carlos Mullov, Ngoc-Quan Pham, Thai Binh Nguyen, Fabian Retkowski, Stefan Constantin, Enes Ugan, Danni Liu, Zhaolin Li, Sai Koneru, Jan Niehues, Alexander Waibel
| Challenge: | a framework to evaluate low-latency speech translations is currently only limited to specific aspects and is not able to compare different approaches. |
| Approach: | They propose a framework to perform and evaluate low-latency speech translation in realistic conditions. |
| Outcome: | The proposed framework evaluates various aspects of low-latency speech translation under realistic conditions. |
Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach (2024.emnlp-main)
Copied to clipboard
| Challenge: | Generating speech through a pipeline that operates at the text level typically loses nuances, intonations, and non-verbal vocalizations. |
| Approach: | They show that fine-tuning speech representation models on phoneme classification leads to more context-invariant representations, and language models trained on these units achieve comparable lexical comprehension to ones trained on hundred times more data. |
| Outcome: | Recent advances in speech representation modeling have shown that learning language directly from speech is feasible. |
Thesis Proposal: Development of End-to-End Speech Translation Models for Indian Languages (2026.eacl-srw)
Copied to clipboard
| Challenge: | Existing approaches to speech-to-speech translation rely on cascaded pipelines . current approaches rely only on text representations, but they suffer from errors and latency . a new direct speech translation framework is proposed to bridge linguistic gaps . |
| Approach: | They propose a sequence-to-sequence direct speech translation framework that can translate speech from one Indian language to another without relying on intermediate text representations. |
| Outcome: | The proposed framework can translate speech from one Indian language to another without relying on intermediate text representations. |
Worse WER, but Better BLEU? Leveraging Word Embedding as Intermediate in Multitask End-to-End Speech Translation (2020.acl-main)
Copied to clipboard
| Challenge: | Existing studies show that multitask learning improves speech translation performance by utilizing word embedding as the intermediate. |
| Approach: | They propose to use word embedding as an intermediate to improve multitask ST models by utilizing word embeds as input. |
| Outcome: | The proposed model outperforms existing models with sufficient training data but is still lacking in the low-resource scenario. |
Character-level Representations Improve DRS-based Semantic Parsing Even in the Age of BERT (2020.emnlp-main)
Copied to clipboard
| Challenge: | a new method of analysis based on semantic tags demonstrates that character-level representations improve performance across a subset of selected semantic phenomena. |
| Approach: | They combine character-level and contextual language model representations to improve performance on Discourse Representation Structure parsing. |
| Outcome: | The proposed model improves performance on a subset of selected semantic phenomena. |
Duplex Diffusion Models Improve Speech-to-Speech Translation (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to speech-to-speech translation train two separate models or a multitask-learned model with low efficiency and inferior performance. |
| Approach: | They propose a duplex diffusion model that applies diffusion probabilistic models to both sides of a reversible duplex Conformer and enables reverse speech translation by simply flipping the input and output ends. |
| Outcome: | The proposed model achieves the first success of reversible speech translation with significant improvements of ASR-BLEU scores compared with a list of state-of-the-art baselines. |