Papers by Alexandre Mourachko
Aligning Speech Segments Beyond Pure Semantics (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing speech-to-speech parallel data is scarce and expensive to create from scratch. |
| Approach: | They propose an algorithm which automatically aligns pairs of speech segments aligned in meaning and expressivity. |
| Outcome: | The proposed algorithm outperforms semantic-focused approaches on content translation quality. |
BOUQuET : dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation (2025.emnlp-main)
Copied to clipboard
Pierre Andrews, Mikel Artetxe, Mariano Coria Meglioli, Marta R. Costa-jussà, Joe Chuang, David Dale, Mark Duppenthaler, Nathanial Paul Ekberg, Cynthia Gao, Daniel Edward Licht, Jean Maillard, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Eduardo Sánchez, Ioannis Tsiamas, Arina Turkatenko, Albert Ventayol-Boada, Shireen Yates
| Challenge: | BOUQUET is a multi-way, multicentric and multi-register/domain dataset and benchmark . the dataset is handcrafted in 8 non-English languages . |
| Approach: | They propose to use BOUQuET to collect a multi-way, multicentric and multi-register/domain dataset and benchmark in 8 non-English languages. |
| Outcome: | The proposed dataset is available at https://huggingface.co/datasets/facebook/bouquet. |
xSIM++: An Improved Proxy to Bitext Mining Performance for Low-Resource Languages (2023.acl-short)
Copied to clipboard
| Challenge: | xsim++ provides a reliable proxy for bitext mining without expensive pipelines. |
| Approach: | They propose a new proxy proxy based on similarity in a multilingual embedding space . they validate this proxy by running a significant number of bitext mining experiments for a set of low-resource languages and then train NMT systems on the mined data. |
| Outcome: | The proposed proxy improves on xsim++ and trains on the mined data. |
LCFO: Long Context and Long Form Output Dataset and Benchmarking (2025.findings-acl)
Copied to clipboard
Marta R. Costa-jussà, Pierre Andrews, Mariano Coria Meglioli, Joy Chen, Joe Chuang, David Dale, Christophe Ropers, Alexandre Mourachko, Eduardo Sánchez, Holger Schwenk, Tuan A. Tran, Arina Turkatenko, Carleigh Wood
| Challenge: | Using long text outputs to evaluate progress in summarization and summary expansion tasks is challenging. |
| Approach: | They propose a framework for assessing gradual summarization and summary expansion capabilities across diverse domains. |
| Outcome: | The proposed framework provides alignments between specific QA pairs and corresponding summaries in 7 domains. |
MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector (2024.findings-acl)
Copied to clipboard
Marta Costa-jussà, Mariano Meglioli, Pierre Andrews, David Dale, Prangthip Hansanti, Elahe Kalbassi, Alexandre Mourachko, Christophe Ropers, Carleigh Wood
| Challenge: | Existing studies on text-based toxicity detection for other languages are limited, especially for languages other than English. |
| Approach: | They propose a multilingual audio-based toxicity classifier which covers 14 different linguistic families and a dataset of 20,000 audio utterances for English and Spanish. |
| Outcome: | The new classifier improves F1-Score by an average of 100% when compared to existing wordlist-based classifiers. |
stopes - Modular Machine Translation Pipelines (2022.emnlp-demos)
Copied to clipboard
Pierre Andrews, Guillaume Wenzek, Kevin Heffernan, Onur Çelebi, Anna Sun, Ammar Kamran, Yingzhe Guo, Alexandre Mourachko, Holger Schwenk, Angela Fan
| Challenge: | Neural machine translation is a natural language deep learning application that needs data to be trained. |
| Approach: | They describe a framework that empowers scalability and versatility for research use cases. |
| Outcome: | The proposed framework empowers scalability and versatility for research use cases. |
BLASER: A Text-Free Speech-to-Speech Translation Evaluation Metric (2023.acl-long)
Copied to clipboard
Mingda Chen, Paul-Ambroise Duquenne, Pierre Andrews, Justine Kao, Alexandre Mourachko, Holger Schwenk, Marta R. Costa-jussà
| Challenge: | End-to-End speech-to speech translation is generally evaluated with text-based metrics . this means generated speech has to be automatically transcribed, making the evaluation dependent on ASR systems. |
| Approach: | They propose a text-free evaluation metric for end-to-end speech-tospeech translation, named BLASER, to avoid the dependency on automatic speech recognition systems. |
| Outcome: | The proposed metric avoids the dependency on automatic speech recognition systems by encoding generated speech segments into a shared embedding space. |