Papers by Ann Lee
SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations (2023.acl-long)
Copied to clipboard
Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswami, Changhan Wang, Juan Pino, Benoît Sagot, Holger Schwenk
| Challenge: | SpeechMatrix is a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. |
| Approach: | They present a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. |
| Outcome: | The proposed model can train bilingual models on 136 language pairs with 418 thousand hours of speech. |
fairseq Sˆ2: A Scalable and Integrable Speech Synthesis Toolkit (2021.emnlp-demo)
Copied to clipboard
| Challenge: | Speech synthesis is the task of generating speech waveforms with desired characteristics, including but not limited to textual content, speaker identity, and speaking styles. |
| Approach: | They propose a fairseq extension for speech synthesis that implements autoregressive and non-AR text-to-speech models and their multi-speaker variants. |
| Outcome: | The proposed extension can train autoregressive and non-AR models and their multi-speaker variants with less curated data and has automatic metrics to facilitate faster iteration and analysis. |
Text-Free Prosody-Aware Generative Spoken Language Modeling (2022.acl-long)
Copied to clipboard
Eugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi, Jade Copet, Kushal Lakhotia, Tu Anh Nguyen, Morgane Riviere, Abdelrahman Mohamed, Emmanuel Dupoux, Wei-Ning Hsu
| Challenge: | Experimental results show that generative spoken language models (LMs) are natural unsupervised multitask learners. |
| Approach: | They propose a prosody-aware generative spoken language model that uses discovered units to generate natural, meaningful, and coherent speech. |
| Outcome: | The proposed model can generate natural, meaningful, and coherent speech given a spoken prompt. |
Speech-to-Speech Translation for a Real-world Unwritten Language (2023.findings-acl)
Copied to clipboard
Peng-Jen Chen, Kevin Tran, Yilin Yang, Jingfei Du, Justine Kao, Yu-An Chung, Paden Tomasello, Paul-Ambroise Duquenne, Holger Schwenk, Hongyu Gong, Hirofumi Inaguma, Sravya Popuri, Changhan Wang, Juan Pino, Wei-Ning Hsu, Ann Lee
| Challenge: | a new study examines speech-to-speech translation (S2ST) that translates speech from one language into another . the research area for unwritten languages remains a research area with little exploration due to the lack of training data. |
| Approach: | They propose a system that translates speech from one language into another . they use Taiwanese Hokkien as an example of an unwritten language . |
| Outcome: | The proposed system can be used to train models in languages without standard writing systems. |
UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units (2023.acl-long)
Copied to clipboard
Hirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen, Changhan Wang, Yu-An Chung, Yun Tang, Ann Lee, Shinji Watanabe, Juan Pino
| Challenge: | Experimental evaluations show that UnitY outperforms a single-pass speech-to-unit translation model by 2.5-4.2 ASR-BLEU with 2.83x decoding speed-up. |
| Approach: | They propose a two-pass direct S2ST architecture which generates textual representations and predicts discrete acoustic units . they show that UnitY outperforms a single-pass speech-to-unit translation model by 2.5-4.2 ASR-BLEU with 2.83x decoding speed-up. |
| Outcome: | The proposed architecture outperforms a single-pass speech-to-unit translation model by 2.5-4.2 ASR-BLEU with 2.83x decoding speed-up on large datasets. |
VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation (2021.acl-long)
Copied to clipboard
Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, Emmanuel Dupoux
| Challenge: | VoxPopuli provides 400K hours of unlabeled speech data in 23 languages . large amounts of multilingual audio data are needed to achieve similar progress for multilingual ASR and ST. |
| Approach: | They propose a large-scale multilingual corpus that provides 400K hours of unlabeled speech data in 23 languages. |
| Outcome: | The proposed corpus provides 400K hours of unlabeled speech data in 23 languages and 1.8K hours transcribed speeches in 15 languages and their aligned oral interpretations into 15 target languages totaling 17.3K hours. |
Textless Speech-to-Speech Translation on Real Data (2022.naacl-main)
Copied to clipboard
Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Yossi Adi, Juan Pino, Jiatao Gu, Wei-Ning Hsu
| Challenge: | Existing text-based speech-to-speech translation systems rely on cascaded approach . text-to text translation systems require text generation and a single input to generate output . |
| Approach: | They propose a textless speech-to-speech translation system that can translate speech from one language into another without the need of text data. |
| Outcome: | The proposed system can translate speech from one language into another without text data. |
Direct Speech-to-Speech Translation With Discrete Units (2022.acl-long)
Copied to clipboard
Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Sravya Popuri, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, Juan Pino, Wei-Ning Hsu
| Challenge: | Existing direct speech-to-speech translation models rely on text generation as an intermediate step. |
| Approach: | They propose a direct speech-to-speech translation model that translates speech from one language to another without relying on intermediate text generation. |
| Outcome: | The proposed model produces 6.7 BLEUs in the Fisher Spanish-English dataset when trained without any text transcripts and with text supervision. |
textless-lib: a Library for Textless Spoken Language Processing (2022.naacl-demo)
Copied to clipboard
Eugene Kharitonov, Jade Copet, Kushal Lakhotia, Tu Anh Nguyen, Paden Tomasello, Ann Lee, Ali Elkahky, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, Yossi Adi
| Challenge: | Textless spoken language processing is an exciting area of research that promises to extend applicability of the standard NLP toolset onto spoken language and languages with few or no textual resources. |
| Approach: | They introduce textless-lib, a PyTorch-based library that provides textless spoken language processing tools. |
| Outcome: | The proposed library significantly simplifies research in the textless setting and will be a handful for speech researchers and the NLP community at large. |
A Comparative Multidimensional Analysis of Empathetic Systems (2024.eacl-long)
Copied to clipboard
| Challenge: | Empathetic dialogue systems have received significant attention, but no systematic review has verified these limitations. |
| Approach: | They analyze 21 empathetic dialogue systems using automated methods to examine their progress. |
| Outcome: | The results show that empathetic dialogue systems lack specificity, reflection levels, diversity . the results also offer guidance for developing future systems . |
Discriminative Reranking for Neural Machine Translation (2021.acl-long)
Copied to clipboard
| Challenge: | reranking models allow the integration of rich features to select a better output hypothesis within an n-best list or lattice. |
| Approach: | They use discriminative reranking to train a large transformer architecture to train an ranked list of hypotheses. |
| Outcome: | Experiments on four WMT directions show that discriminative reranking improves translation quality. |
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent expressive speech-to-speech translation systems have achieved impressive expressivity preservation performances by cascading unit-to speech (U2S) generator to the speech- to-unit translation model. |
| Approach: | They propose a textless acoustic model with a self-supervised distillation strategy for noise-robust expressive speech-to-speech translation (S2ST) They aim to address this limitation by incorporating a distillation with no label (DINO) self-controlled training strategy into the model’s pretraining process. |
| Outcome: | The proposed model significantly improved the expressive speech-to-speech translation system in noisy environments while maintaining competitive performance in clean environments. |