Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for detecting AI-generated music are weak and vulnerable to audio perturbations. |
| Approach: | They propose a multimodal late-fusion pipeline that combines automatically transcribed sung lyrics and speech features capturing lyrics related information within the audio. |
| Outcome: | The proposed method outperforms existing detectors while being more robust to audio perturbations. |
Similar Papers
Unsupervised Melody-to-Lyrics Generation (2023.acl-long)
Copied to clipboard
Yufei Tian, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone, Gunnar Sigurdsson, Chenyang Tao, Wenbo Zhao, Tagyoung Chung, Jing Huang, Nanyun Peng
| Challenge: | Existing methods for automatic melody-to-lyric generation are limited due to the limited amount of melody-lyrical aligned data. |
| Approach: | They propose a method for automatic melody-to-lyric generation without training on any aligned melody-lyr data. |
| Outcome: | The proposed model generates high-quality lyrics that are singable, intelligible, and coherent than baseline models. |
A Melody-Conditioned Lyrics Language Model (N18-1)
Copied to clipboard
Kento Watanabe, Yuichiroh Matsubayashi, Satoru Fukayama, Masataka Goto, Kentaro Inui, Tomoyasu Nakano
| Challenge: | Existing models for lyrics generation are insufficient to capture relationship between lyrics and melody. |
| Approach: | They propose a data-driven language model that generates entire lyrics for a given melody. |
| Outcome: | The proposed model generates fluent lyrics while maintaining compatibility between lyrics and melodies. |
SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition (2025.acl-long)
Copied to clipboard
Shuangrui Ding, Zihan Liu, Xiaoyi Dong, Pan Zhang, Rui Qian, Junhao Huang, Conghui He, Dahua Lin, Jiaqi Wang
| Challenge: | Creating lyrics and melodies in symbolic format requires expert knowledge of melody and an advanced understanding of lyrics. |
| Approach: | They introduce SongComposer, a music-specialized large language model that can create symbolic lyrics and melodies following instructions. |
| Outcome: | The proposed model outperforms existing models in symbolic song composition tasks. |
Lyrics Segmentation: Textual Macrostructure Detection using Convolutions (C18-1)
Copied to clipboard
| Challenge: | Lyrics contain repeated patterns that are correlated with the repetitions found in the music they accompany. |
| Approach: | They propose to apply a convolutional neural network to a task to detect lyrics by using a neural network. |
| Outcome: | The proposed features improve the state-of-the-art in lyrics segmentation . a convolutional neural network is applied to the task and it is able to detect lyrics in different genres. |
ALCAP: Alignment-Augmented Music Captioner (2023.emnlp-main)
Copied to clipboard
| Challenge: | Traditional approaches to music captioning ignore the intricate interplay between the two . however, a comprehensive understanding of music necessitates the integration of both these elements. |
| Approach: | They propose a method to learn multimodal alignment between audio and lyrics through contrastive learning. |
| Outcome: | The proposed method achieves new state-of-the-art on two music captioning datasets. |
Syllable-level lyrics generation from melody exploiting character-level language model (2024.findings-eacl)
Copied to clipboard
| Challenge: | Pre-trained language models specifically designed at the syllable level are not available. |
| Approach: | They propose to exploit character-level language models for syllable-level lyrics generation from symbolic melody. |
| Outcome: | The proposed system improves coherence and correctness of generated lyrics without training expensive language models. |
UniLG: A Unified Structure-aware Framework for Lyrics Generation (2023.acl-long)
Copied to clipboard
| Challenge: | Existing works ignore musical attributes hidden behind lyrics and structure of lyrics . existing works ignore structure of generated lyrics and do not consider structure of songs . |
| Approach: | They propose a framework for conditional lyrics generation that considers structure and relationship between lyrics and music. |
| Outcome: | The proposed framework improves the structure modeling and unifies different conditions for different types of lyrics generation. |
Reasoning-Aware AIGC Detection via Alignment and Reinforcement (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to AIGC detection have relied on statistical classifiers or black-box neural models, which exploit surface-level patterns and struggle to generalize as LLMs evolve. |
| Approach: | They propose a framework that generates interpretable reasoning chains before classification using supervised fine-tuning and reinforcement learning to improve accuracy. |
| Outcome: | The proposed framework achieves state-of-the-art performance across multiple benchmarks, offering a robust and transparent solution for AIGC detection. |
Robust AI-Generated Text Detection by Restricted Embeddings (2024.findings-emnlp)
Copied to clipboard
Kristian Kuznetsov, Eduard Tulchinskii, Laida Kushnareva, German Magai, Serguei Barannikov, Sergey Nikolenko, Irina Piontkovskaya
| Challenge: | Existing approaches for artificial text detection are score-based and classifier-based . however, score-driven methods often rely on a score-derived score. |
| Approach: | They investigate the ability of classifier-based detectors to transfer to unseen generators or semantic domains. |
| Outcome: | The proposed methods improve the out-of-distribution classification score by up to 9% and 14%. |
A Unified Feature Mixture Framework for Joint Speech and Singing Deepfake Detection (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for deepfake detection fail under speech-to-singing domain shift . a speech-retentive multi-domain fine-tuning strategy enables adaptation to singing . |
| Approach: | They propose a unified deepfake detector based on a multi-branch mixture-of-experts architecture that integrates three complementary feature views. |
| Outcome: | The proposed detector achieves 1.82% EER on CtrSVDD, compared to 37–62% for existing detectors . it can generalize to unseen generators and preserve strong speech performance . |