Challenge: Large Language Models have shown immense potential in multimodal applications, but convergence between textual and musical domains remains unexplored.
Approach: They propose a system that aligns music representations with a frozen LLM . they train the system on an extensive music caption dataset and fine-tune it with instructional data .
Outcome: The proposed system bridges the gap between music audio and textual contexts by combining music captions with a frozen model . it performs well in generating music caption and composing music-related Q&A pairs . the proposed system is available for free download at http://www.musilingo.com/ .

Similar Papers

Mustango: Toward Controllable Text-to-Music Generation (2024.naacl-long)

Copied to clipboard

Challenge: Mustango is a text-to-music system that allows music-domain-knowledge-informed text-based music generation.
Approach: They propose a music-domain-knowledge-inspired text-to-music system based on diffusion that generates music with captions that include specific instructions related to chords, beats, key and tempo.
Outcome: The proposed system outperforms existing models in music generation tasks.
CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Current music information retrieval systems struggle to meet linguistic diversity challenges . current systems struggle with text queries in non-English languages .
Approach: They propose a music information retrieval system that supports both ABC notation and MIDI . CLaMP 2 includes a multilingual text encoder and a multiple-modal music encoder .
Outcome: The proposed system achieves state-of-the-art results in multilingual semantic search and music classification across modalities.
DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in music large language models have significantly improved music understanding tasks, but the potential of incorporating additional modalities such as images, videos and textual music features remains unexplored.
Approach: They propose a multimodal music understanding LLM fine-tuned via multi-way instruction tuning with multi-ways aligned music, text, image, and video data.
Outcome: The proposed model achieves state-of-the-art performance across six music understanding tasks and zero-shot scenarios.
FIGMA: Towards FIne-Grained Music retrievAl (2026.acl-long)

Copied to clipboard

Challenge: Existing music retrieval models fail to retrieve fine-grained musical attributes when using coarse semantic queries.
Approach: They propose a multi-view contrastive architecture that captures high-level semantic context and fine-grained musical attributes within a unified representation space.
Outcome: The proposed method outperforms existing CLAP-based music retrieval models on multiple benchmarks.
Thesis Proposal: Multimodal Benchmark for Music Understanding in Large Language Models (2026.eacl-srw)

Copied to clipboard

Challenge: Existing music-focused benchmarks are fragmented, largely single-modality, Western-centric . existing methods for evaluating MLLMs are lacking reproducibility and reliability .
Approach: They propose to develop a musically multimodal benchmark that will integrate music into the benchmark.
Outcome: The proposed benchmark will integrate culturally diverse musical material beyond the dominant Western canon.
CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages (2025.findings-acl)

Copied to clipboard

Challenge: Music information retrieval (MIR) is a field that aims at developing computational tools for processing, organizing, and accessing music data.
Approach: They propose a framework that aligns music modalities with multilingual text in a shared representation space.
Outcome: Experiments show CLaMP 3 performs state-of-the-art on multiple MIR tasks . it surpasses baselines and shows excellent generalization in multimodal and multilingual contexts .
ALCAP: Alignment-Augmented Music Captioner (2023.emnlp-main)

Copied to clipboard

Challenge: Traditional approaches to music captioning ignore the intricate interplay between the two . however, a comprehensive understanding of music necessitates the integration of both these elements.
Approach: They propose a method to learn multimodal alignment between audio and lyrics through contrastive learning.
Outcome: The proposed method achieves new state-of-the-art on two music captioning datasets.
Moûsai: Efficient Text-to-Music Diffusion Models (2024.acl-long)

Copied to clipboard

Challenge: Recent years have seen the rapid development of large generative models for text; however, little research has explored the connection between text and another “language” of communication – music.
Approach: They develop a text-to-music generation model that can generate multiple minutes of high-quality stereo music at 48kHz from textual descriptions.
Outcome: The proposed model can generate multiple minutes of high-quality stereo music at 48kHz from textual descriptions.
Generative Music Models’ Alignment with Professional and Amateur Users’ Expectations (2025.findings-acl)

Copied to clipboard

Challenge: Recent years have witnessed rapid advances in text-to-music generation using large language models.
Approach: They propose a task to align AI-generated music with human expressions . they use a dataset of over 1.5 million songs to analyze their content .
Outcome: The proposed framework outperforms baseline models and facilitates end-to-end generation of songs audio.
Learning Musical Representations for Music Performance Question Answering (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for audio-visual learning fail to consider the distinctive characteristics of instruments and music.
Approach: They propose to integrate multimodal interactions within the context of music data and annotate and release rhythmic and music sources in the current music datasets to enable the model to learn music characteristics.
Outcome: The proposed model can learn music characteristics from the current music datasets and align its predictions with the temporal dimension.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations