Papers by Yossi Adi

15 papers
Slamming: Training a Speech Language Model on One GPU in a Day (2025.findings-acl)

Copied to clipboard

Challenge: *Slam* is a recipe for training high-quality Speech Language Models (SLMs) on a single academic GPU in 24 hours.
Approach: They propose a recipe for training high-quality Speech Language Models on a single academic GPU in 24 hours.
Outcome: The proposed training recipe outperforms predicted compute optimal performance, giving an optimistic view to SLM feasibility.
fairseq Sˆ2: A Scalable and Integrable Speech Synthesis Toolkit (2021.emnlp-demo)

Copied to clipboard

Challenge: Speech synthesis is the task of generating speech waveforms with desired characteristics, including but not limited to textual content, speaker identity, and speaking styles.
Approach: They propose a fairseq extension for speech synthesis that implements autoregressive and non-AR text-to-speech models and their multi-speaker variants.
Outcome: The proposed extension can train autoregressive and non-AR models and their multi-speaker variants with less curated data and has automatic metrics to facilitate faster iteration and analysis.
Text-Free Prosody-Aware Generative Spoken Language Modeling (2022.acl-long)

Copied to clipboard

Challenge: Experimental results show that generative spoken language models (LMs) are natural unsupervised multitask learners.
Approach: They propose a prosody-aware generative spoken language model that uses discovered units to generate natural, meaningful, and coherent speech.
Outcome: The proposed model can generate natural, meaningful, and coherent speech given a spoken prompt.
Textless Speech-to-Speech Translation on Real Data (2022.naacl-main)

Copied to clipboard

Challenge: Existing text-based speech-to-speech translation systems rely on cascaded approach . text-to text translation systems require text generation and a single input to generate output .
Approach: They propose a textless speech-to-speech translation system that can translate speech from one language into another without the need of text data.
Outcome: The proposed system can translate speech from one language into another without text data.
Direct Speech-to-Speech Translation With Discrete Units (2022.acl-long)

Copied to clipboard

Challenge: Existing direct speech-to-speech translation models rely on text generation as an intermediate step.
Approach: They propose a direct speech-to-speech translation model that translates speech from one language to another without relying on intermediate text generation.
Outcome: The proposed model produces 6.7 BLEUs in the Fisher Spanish-English dataset when trained without any text transcripts and with text supervision.
Speaking Style Conversion in the Waveform Domain Using Discrete Self-Supervised Units (2023.findings-emnlp)

Copied to clipboard

Challenge: DISSC is a lightweight voice conversion method that converts the rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner.
Approach: They propose a method that converts rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner.
Outcome: The proposed method outperforms baseline methods on quantitative and qualitative evaluations.
textless-lib: a Library for Textless Spoken Language Processing (2022.naacl-demo)

Copied to clipboard

Challenge: Textless spoken language processing is an exciting area of research that promises to extend applicability of the standard NLP toolset onto spoken language and languages with few or no textual resources.
Approach: They introduce textless-lib, a PyTorch-based library that provides textless spoken language processing tools.
Outcome: The proposed library significantly simplifies research in the textless setting and will be a handful for speech researchers and the NLP community at large.
Generative Spoken Dialogue Language Modeling (2023.tacl-1)

Copied to clipboard

Challenge: dGSLM is the first “textless” model able to generate audio samples of naturalistic spoken dialogues.
Approach: They propose a model that generates speech, laughter, and other paralinguistic signals in two channels simultaneously and reproduces more naturalistic turn taking compared to a text-based cascaded model.
Outcome: The proposed model reproduces more naturalistic and fluid turn taking than a text-based cascaded model.
Generative Spoken Language Model based on continuous word-sized audio tokens (2023.emnlp-main)

Copied to clipboard

Challenge: Text-based language models outperform character-based models, but speech inputs are 20ms or 40ms-long discrete units.
Approach: They propose a generative language model based on word-size continuous audio tokens . they replace lookup table for lexical types with a Lexical Embedding function .
Outcome: The proposed model is five times more memory efficient than discrete unit GSLMs and is phonetically and semantically interpretable.
Transformers are Multi-State RNNs (2024.emnlp-main)

Copied to clipboard

Challenge: Schwartz et al., 2017) have been using transformers for long-range tasks for NLP since the 1990s.
Approach: They propose a transformer-only transformer with unlimited hidden state size that can be converted into bounded multistate RNNs by fixing the size of their hidden state.
Outcome: The proposed compression policy outperforms baseline compression policies on long range tasks and LLMs.
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion (2026.acl-short)

Copied to clipboard

Challenge: Large Language Models lack visual grounding on visual reasoning, despite training on text alone.
Approach: They propose a late multi-image fusion method that augments LLMs with test-time visual signals.
Outcome: Using a late multi-image fusion method, the proposed model outperforms LLMs on visual reasoning and matches VLMs in vision-based tasks.
GmSLM : Generative Marmoset Spoken Language Modeling (2025.findings-emnlp)

Copied to clipboard

Challenge: Generative Marmoset Spoken Language Modeling (GmSLM) is an optimized spoken language model pipeline for Marmosaet vocal communication.
Approach: They propose an optimized spoken language model pipeline for Marmoset vocal communication.
Outcome: The proposed framework can be used to link vocal communication with brain activity in Marmoset monkeys.
StressTest: Can YOUR Speech LM Handle the Stress? (2026.findings-acl)

Copied to clipboard

Challenge: Recent speech-aware language models (SLMs) have enabled direct audio processing, allowing models to access the full expressive range of spoken language.
Approach: They propose a data generation pipeline that simulates change of meaning implied by stress variation and propose 'stresstest' to evaluate models' ability to distinguish between meanings of speech based on stress pattern.
Outcome: The proposed model outperforms existing models on sentence stress reasoning and detection.
Textless Speech Emotion Conversion using Discrete & Decomposed Representations (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for modifying emotion of speech are difficult because emotion affects all levels simultaneously.
Approach: They propose a method to convert a spoken language speech into a model of emotion . they use phonetic-content units, prosodic features, speaker, and emotion to modify the emotion a speech utterance has.
Outcome: The proposed method beats text-based systems in terms of perceived emotion and audio quality.
On Generative Spoken Language Modeling from Raw Audio (2021.tacl-1)

Copied to clipboard

Challenge: Using a set of metrics to evaluate the learned representations, we aim to create a system that learns from natural interactions as infants learn their first language.
Approach: They propose a task of learning acoustic and linguistic characteristics from raw audio and a set of metrics to evaluate the learned representations at acustic, linguistic and encoding levels.
Outcome: The proposed models evaluate the learned representations at acoustic and linguistic levels for both encoding and generation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations