Moûsai: Efficient Text-to-Music Diffusion Models (2024.acl-long)

Copied to clipboard

Challenge: Recent years have seen the rapid development of large generative models for text; however, little research has explored the connection between text and another “language” of communication – music.
Approach: They develop a text-to-music generation model that can generate multiple minutes of high-quality stereo music at 48kHz from textual descriptions.
Outcome: The proposed model can generate multiple minutes of high-quality stereo music at 48kHz from textual descriptions.

Similar Papers

Mustango: Toward Controllable Text-to-Music Generation (2024.naacl-long)

Copied to clipboard

Challenge: Mustango is a text-to-music system that allows music-domain-knowledge-informed text-based music generation.
Approach: They propose a music-domain-knowledge-inspired text-to-music system based on diffusion that generates music with captions that include specific instructions related to chords, beats, key and tempo.
Outcome: The proposed system outperforms existing models in music generation tasks.
Rapid Diffusion: Building Domain-Specific Text-to-Image Synthesizers with Fast Inference Speed (2023.acl-industry)

Copied to clipboard

Challenge: Text-to-Image Synthesis (TIS) aims to generate images based on textual inputs . but, current diffusion-based models lack entity knowledge and low inference speed .
Approach: They propose a framework for training and deploying latent diffusion models with rich entity knowledge injected and optimized networks.
Outcome: The proposed framework improves image quality and inference speed and can be used in industrial applications.
Kandinsky 3: Text-to-Image Synthesis for Multifunctional Generative Framework (2024.emnlp-demo)

Copied to clipboard

Challenge: Text-to-image (T2I) diffusion models are popular for image manipulation, but also for video generation.
Approach: They propose a novel T2I diffusion model based on latent diffusion that extends the base model for various applications.
Outcome: The proposed model achieves high quality and photorealism and is 3 times faster than the base model.
MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response (2024.findings-naacl)

Copied to clipboard

Challenge: Large Language Models have shown immense potential in multimodal applications, but convergence between textual and musical domains remains unexplored.
Approach: They propose a system that aligns music representations with a frozen LLM . they train the system on an extensive music caption dataset and fine-tune it with instructional data .
Outcome: The proposed system bridges the gap between music audio and textual contexts by combining music captions with a frozen model . it performs well in generating music caption and composing music-related Q&A pairs . the proposed system is available for free download at http://www.musilingo.com/ .
Leveraging Pretrained Language Models as Energy Functions for Glauber Dynamics Text Diffusion (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for generating text using Glauber dynamics are autoregressive, but they face a number of limitations.
Approach: They propose a discrete diffusion-based generative model for text generation using Glauber dynamics from statistical physics and use pretrained causal/masked language models to improve the quality of the generated text.
Outcome: The proposed model outperforms existing models on some common sense reasoning tasks and planning/search tasks.
ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling (2026.acl-long)

Copied to clipboard

Challenge: Recent efforts on text-to-audio generation are exploring fine-grained controllability . however, their performance at scale is limited due to data scarcity .
Approach: They propose a multi-task learning problem for high-controllability text-to-audio generation . they propose scalable diffusion transformers that augment condition information in sequence .
Outcome: The proposed method outperforms existing methods on objective and subjective evaluations.
Music for All: Representational Bias and Cross-Cultural Adaptability of Music Generation Models (2025.findings-naacl)

Copied to clipboard

Challenge: Existing music generation models are limited in their coverage of the musical genres and cultures of the world.
Approach: They propose to use parametric fine tuning techniques to mitigat the bias in existing music datasets.
Outcome: The proposed models are able to perform well across genres and cultures.
Can Diffusion Model Achieve Better Performance in Text Generation ? Bridging the Gap between Training and Inference ! (2023.findings-acl)

Copied to clipboard

Challenge: Existing models for text generation use a discrete data embedding module to map the data into the continuous space.
Approach: They propose two methods to bridge the gap between training and inference by mapping the discrete text into the continuous space.
Outcome: The proposed methods can achieve 100 200 speedup with better performance on 6 generation tasks.
Segment-Level Diffusion: A Framework for Controllable Long-Form Generation with Diffusion Language Models (2025.acl-long)

Copied to clipboard

Challenge: Diffusion models have shown promise in text generation, but often struggle with generating long, coherent, and contextually accurate text.
Approach: They propose a framework that enhances diffusion-based text generation through text segmentation, robust representation training with adversarial and contrastive learning, and improved latent-space guidance.
Outcome: The proposed framework improves diffusion-based text generation and improves scalability and fluency.
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions (2026.acl-long)

Copied to clipboard

Challenge: Generative audio modeling has been fragmented into specialized tasks such as text-to-speech (TTS), text- to-music (TTM), and text-ta (TTA) specialized models require reference audio for timbre cloning and strict phoneme alignment, whereas TTA models generate unstructured textures from open-ended captions.
Approach: They propose a unified flow-matching framework capable of synthesizing speech, music, sound effects . they propose 'token injection mechanism' that projects unstructured environmental sounds into structured temporal latent space .
Outcome: The proposed framework achieves state-of-the-art performance in instruction-based TTS and TTM while maintaining competitive fidelity in TTA.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations