Challenge: Existing methods for expressive text-to-speech only implicitly learn prosody with masked token reconstruction tasks.
Approach: They propose a cross-modal contrastive pre-training framework that learns from prosody variance of the same text token under different contexts.
Outcome: The proposed framework can learn from prosody variance of a text token under different contexts.

Similar Papers

Cross-modal Contrastive Learning for Speech Translation (2022.naacl-main)

Copied to clipboard

Challenge: Existing approaches for speech translation focus on using additional data from MT and automatic speech recognition (ASR).
Approach: They propose a cross-modal contrastive learning method for end-to-end speech-totext translation.
Outcome: The proposed method outperforms existing methods on a popular benchmark MuST-C.
Prosody-TTS: Improving Prosody with Masked Autoencoder and Conditional Diffusion Model For Expressive Text-to-Speech (2023.findings-acl)

Copied to clipboard

Challenge: Expressive text-to-speech aims to generate high-quality samples with rich prosody . prosodic attributes in highly dynamic voices are difficult to capture and model without intonation .
Approach: They propose a pipeline that enhances prosody modeling and sampling by introducing a self-supervised masked autoencoder and a diffusion model to sample diverse prosodic patterns within the latent space.
Outcome: The proposed pipeline achieves new state-of-the-art in text-to-speech with natural and expressive synthesis.
On the Language Encoder of Contrastive Cross-modal Models (2024.findings-acl)

Copied to clipboard

Challenge: Pretrained audio-language models such as AudioCLIP and AudioCLAP have shown promising results on vision-language (VL) tasks.
Approach: They extensively evaluate how unsupervised and supervised sentence embedding training affect language encoder quality and cross-modal task performance.
Outcome: The proposed model improves on visual-language (VL) and audio-language tasks when the amount of training data is large.
CLASP: Cross-modal Alignment Using Pre-trained Unimodal Models (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in speech-text pretraining rely on parallel speech- text data . however, data accessibility is a challenge due to the limited data available.
Approach: They propose a framework for jointly performing speech and text processing without parallel corpora during pre-training but only downstream.
Outcome: The proposed framework extracts distinct representations for speech and text, aligning them effectively in a newly defined space using a multi-level contrastive learning mechanism.
Do Audio-Language Models Understand Linguistic Variations? (2025.naacl-short)

Copied to clipboard

Challenge: Existing open-vocabulary audio language models struggle to generalize to linguistic variations in textual queries.
Approach: They propose a novel technique to learn audio-language representations agnostic to linguistic variations by reformulating contrastive loss used in CLAP architectures.
Outcome: The proposed approach improves the performance of the open-vocabulary audio language models by 0.8%-13% across benchmarks and enhances robustness to linguistic variation.
SoftMCL: Soft Momentum Contrastive Learning for Fine-grained Sentiment-aware Pre-training (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for pre-training language models capture general language understanding but fail to distinguish affective impact of a particular context to a specific word.
Approach: They propose a soft momentum contrastive learning method for fine-grained sentiment-aware pre-training that uses valence ratings as soft-label supervision instead of hard labels.
Outcome: The proposed method improves on four sentiment-related tasks and the results are published online.
CTAL: Pre-training Cross-modal Transformer for Audio-and-Language Representations (2021.emnlp-main)

Copied to clipboard

Challenge: Existing audio-language task-specific predictive approaches focus on building complicated late-fusion mechanisms.
Approach: They propose a cross-modal transformer for audio-and-language that learns inter-modal connections between audio and language through two proxy tasks on a large amount of audio- and-language pairs.
Outcome: The proposed model improves on multiple audio-and-language tasks and can be used in fine-tuning phase.
ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models (2025.coling-main)

Copied to clipboard

Challenge: Text-to-speech (TTS) models have been developed to generate high-quality speech.
Approach: They propose an end-to-end TTS model that integrates large self-supervised speech models and conditional flow matching to model prosodic features effectively.
Outcome: The proposed model improves synthesis quality and efficiency compared to existing models, showing that it generates more prosodic and expressive speech synthesizing.
FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining (2026.acl-long)

Copied to clipboard

Challenge: Existing audio-language models excel at clip-level understanding but struggle with frame-level tasks.
Approach: They propose a novel training paradigm that advances both clip- and frame-level alignment in CLAP with heterogeneous data.
Outcome: The proposed training paradigm improves both clip- and frame-level alignment in CLAP with heterogeneous data.
RC3: Regularized Contrastive Cross-lingual Cross-modal Pre-training (2023.findings-acl)

Copied to clipboard

Challenge: Existing V&L pre-training methods rely on strictly-aligned multilingual image-text pairs generated from English-centric datasets.
Approach: They propose a regularized cross-lingual visual contrastive learning objective that constrains representation proximity of weakly-aligned multilingual image-text pairs.
Outcome: The proposed model outperforms competing models with weak zero-shot capability on 5 multi-modal tasks across 6 languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations