Papers by Emmanuel Dupoux

22 papers
BabyCloud, a Technological Platform for Parents and Researchers (L18-1)

Copied to clipboard

Challenge: a platform for capturing, storing and analyzing day-long audio recordings and photos of children's linguistic environments is proposed . the proposed platform connects families and academics, with strong innovation potential for each type of users.
Approach: They propose a platform for capturing, storing and analyzing audio recordings and photos of children's linguistic environments.
Outcome: The proposed platform connects families and academics with strong innovation potential for each type of users.
SpiRit-LM: Interleaved Spoken and Written Language Model (2025.tacl-1)

Copied to clipboard

Challenge: SpiRit-LM is a foundation multimodal language model that freely mixes text and speech.
Approach: They propose a multimodal language model that freely mixes text and speech . they extend the model to the speech modality by continuously training it on text and language units.
Outcome: The proposed model can learn new tasks in a few-shot fashion across modalities.
Word-order Biases in Deep-agent Emergent Communication (P19-1)

Copied to clipboard

Challenge: a recent study examines the "natural" word-order constraints that constrain neural networks . we train models to communicate about paths in a simple gridworld .
Approach: They propose to inoculate a notion of "effort" into neural networks to make their linguistic behavior more human-like.
Outcome: The proposed models show a strong tendency to avoid redundancy and minimize long-distance dependencies.
SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation (2026.acl-long)

Copied to clipboard

Challenge: Empirically, SpidR-Adapt achieves rapid gains in phonemic discriminability and downstream spoken language modeling scores . current self-supervised learning models require thousands of hours of training data to learn meaningful linguistic representations.
Approach: They propose a bi-level optimization framework for rapid adaptation of speech units to new languages using minimal unlabeled data.
Outcome: The proposed model achieves rapid gains in phonemic discriminability and spoken language modeling scores . it surpasses in-domain toplines after training on less than 1h of target-language audio .
Compositionality and Generalization In Emergent Languages (2020.acl-main)

Copied to clipboard

Challenge: a new study examines whether emergent languages possess compositionality . compositionality is a core concept in linguistics, but linguists' definitions assume full knowledge of primitive expressions and their combination rules.
Approach: They propose to use compositionality to combine expressions according to systematic rules to refer to composite concepts.
Outcome: The proposed language has compositionality, but it is not generalized, the authors show . they show that the more compositional a language is, the more easily it will be picked up by new learners .
Identification of Primary and Collateral Tracks in Stuttered Speech (2020.lrec-1)

Copied to clipboard

Challenge: Disfluency detection is a challenging task because of its different metrics depending on whether the input features are text or speech.
Approach: They propose a framework for disfluency detection inspired by the clinical and the natural language processing perspective together with the theory of performance from (Clark, 1998) . they present a forced-aligned disfluence dataset and propose new audio features inspired by word-based span features.
Outcome: The proposed framework outperforms baselines for speech-based predictions on a forced-aligned disfluency dataset from semi-directed interviews.
DP-Parse: Finding Word Boundaries from Raw Speech with an Instance Lexicon (2022.tacl-1)

Copied to clipboard

Challenge: Existing nonparametric models for text segmentation use a Dirichlet process to jointly segment sentences and build a lexicon of word types.
Approach: They propose a Bayesian nonparametric model that uses a Dirichlet process to jointly segment sentences and build a lexicon of word types.
Outcome: The proposed model improves on the Zero Resource Speech Benchmark 2017 and can learn semantic and syntactic representations as assessed by a new spoken word embedding benchmark.
Text-Free Prosody-Aware Generative Spoken Language Modeling (2022.acl-long)

Copied to clipboard

Challenge: Experimental results show that generative spoken language models (LMs) are natural unsupervised multitask learners.
Approach: They propose a prosody-aware generative spoken language model that uses discovered units to generate natural, meaningful, and coherent speech.
Outcome: The proposed model can generate natural, meaningful, and coherent speech given a spoken prompt.
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models (2024.emnlp-main)

Copied to clipboard

Challenge: EmphAssess evaluates speech-to-speech models' ability to encode and reproduce prosodic emphasis across a change of speaker and language.
Approach: They propose a prosodic benchmark to evaluate the ability of speech-to-speech models to encode and reproduce prosodic emphasis.
Outcome: The proposed model can encode and reproduce prosodic emphasis across speech inputs and outputs . EmphaClass classifies emphasis at the frame or word level .
Countering Reward Over-Optimization in LLM with Demonstration-Guided Reinforcement Learning (2024.findings-acl)

Copied to clipboard

Challenge: Existing approaches address ROO by adding KL regularization, requiring computationally expensive hyperparameter tuning.
Approach: They propose a reinforcement learning approach that leverages human demonstrations and a reward model to recalibrate the reward objective.
Outcome: The proposed approach achieves comparable performance to carefully tuned baselines while mitigating ROO in three RL language tasks.
VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation (2021.acl-long)

Copied to clipboard

Challenge: VoxPopuli provides 400K hours of unlabeled speech data in 23 languages . large amounts of multilingual audio data are needed to achieve similar progress for multilingual ASR and ST.
Approach: They propose a large-scale multilingual corpus that provides 400K hours of unlabeled speech data in 23 languages.
Outcome: The proposed corpus provides 400K hours of unlabeled speech data in 23 languages and 1.8K hours transcribed speeches in 15 languages and their aligned oral interpretations into 15 target languages totaling 17.3K hours.
textless-lib: a Library for Textless Spoken Language Processing (2022.naacl-demo)

Copied to clipboard

Challenge: Textless spoken language processing is an exciting area of research that promises to extend applicability of the standard NLP toolset onto spoken language and languages with few or no textual resources.
Approach: They introduce textless-lib, a PyTorch-based library that provides textless spoken language processing tools.
Outcome: The proposed library significantly simplifies research in the textless setting and will be a handful for speech researchers and the NLP community at large.
Generative Spoken Dialogue Language Modeling (2023.tacl-1)

Copied to clipboard

Challenge: dGSLM is the first “textless” model able to generate audio samples of naturalistic spoken dialogues.
Approach: They propose a model that generates speech, laughter, and other paralinguistic signals in two channels simultaneously and reproduces more naturalistic turn taking compared to a text-based cascaded model.
Outcome: The proposed model reproduces more naturalistic and fluid turn taking than a text-based cascaded model.
Generative Spoken Language Model based on continuous word-sized audio tokens (2023.emnlp-main)

Copied to clipboard

Challenge: Text-based language models outperform character-based models, but speech inputs are 20ms or 40ms-long discrete units.
Approach: They propose a generative language model based on word-size continuous audio tokens . they replace lookup table for lexical types with a Lexical Embedding function .
Outcome: The proposed model is five times more memory efficient than discrete unit GSLMs and is phonetically and semantically interpretable.
Frequency & Compositionality in Emergent Communication (2025.emnlp-main)

Copied to clipboard

Challenge: Natural languages exhibit a universal tendency to resist regular patterns, developing idiosyncratic forms.
Approach: They investigate the relationship between frequency and compositionality in emergent languages . they use a referential game setting to manipulate input frequency through Zipfian distributions .
Outcome: The proposed method shows that the frequency distributions of the most frequent words resist regular patterns, resulting in less compositional structure.
MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery (2026.acl-long)

Copied to clipboard

Challenge: MauBERT models learn from multilingual data to predict articulatory features or phones, resulting in language-independent phonetic representations.
Approach: They introduce a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning.
Outcome: The proposed model can predict phonetic features in 55 languages with minimal fine-tuning (10 hours of speech) it is more context-invariant than state-of-the-art models and adapts to unseen languages and casual speech with minimal self-supervised fine- tuning (10 hours)
Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach (2024.emnlp-main)

Copied to clipboard

Challenge: Generating speech through a pipeline that operates at the text level typically loses nuances, intonations, and non-verbal vocalizations.
Approach: They show that fine-tuning speech representation models on phoneme classification leads to more context-invariant representations, and language models trained on these units achieve comparable lexical comprehension to ones trained on hundred times more data.
Outcome: Recent advances in speech representation modeling have shown that learning language directly from speech is feasible.
Seshat: a Tool for Managing and Verifying Annotation Campaigns of Audio Data (2020.lrec-1)

Copied to clipboard

Challenge: Seshat is a software for the automated management of annotation campaigns for audio/speech data.
Approach: They propose a system for the automated management of annotation campaigns for audio/speech data which addresses these challenges.
Outcome: The proposed system computes an associated inter-annotator agreement with the gamma measure taking into account the categorisation and segmentation discrepancies.
XLS-R fine-tuning on noisy word boundaries for unsupervised speech segmentation into words (2023.findings-emnlp)

Copied to clipboard

Challenge: a new method to segment speech into words is needed to overcome the lack of explicit word boundaries in the speech stream.
Approach: They propose to fine-tune a self-supervised speech model to predict word boundaries . they use XLS-R to fine tune the models and infer new word boundary labels .
Outcome: The proposed model outperforms existing models and sets a new state-of-the-art on five corpora with different languages.
Textless Speech Emotion Conversion using Discrete & Decomposed Representations (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for modifying emotion of speech are difficult because emotion affects all levels simultaneously.
Approach: They propose a method to convert a spoken language speech into a model of emotion . they use phonetic-content units, prosodic features, speaker, and emotion to modify the emotion a speech utterance has.
Outcome: The proposed method beats text-based systems in terms of perceived emotion and audio quality.
LongTail-Swap: benchmarking language models’ abilities on rare words (2025.findings-emnlp)

Copied to clipboard

Challenge: LongTail-Swap is a benchmark that focuses on the tail of the word distribution, i.e., measures the ability of LMs to learn new words with very little exposure, like infants do.
Approach: They introduce LongTail-Swap, a benchmark that measures the ability of language models to learn new words with very little exposure, like infants do.
Outcome: The proposed benchmark measures the ability of language models to learn new words with very little exposure, like infants do.
On Generative Spoken Language Modeling from Raw Audio (2021.tacl-1)

Copied to clipboard

Challenge: Using a set of metrics to evaluate the learned representations, we aim to create a system that learns from natural interactions as infants learn their first language.
Approach: They propose a task of learning acoustic and linguistic characteristics from raw audio and a set of metrics to evaluate the learned representations at acustic, linguistic and encoding levels.
Outcome: The proposed models evaluate the learned representations at acoustic and linguistic levels for both encoding and generation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations