Challenge: Experimental results show that multimodal, multimodal approaches to lyrics translation are more effective than text-only approaches.
Approach: They propose a multilingual, multimodal benchmark for singable lyrics translation . they propose syllable-constrained audio-video LLM with Chain-of-Thought .
Outcome: The proposed system outperforms text-based models in singability and contextual accuracy.

Similar Papers

Translate the Beauty in Songs: Jointly Learning to Align Melody and Translate Lyrics (2023.findings-emnlp)

Copied to clipboard

Challenge: Song translation requires both translation of lyrics and alignment of music notes . human translators of songs need to have a mastery of cultural traditions and the poetic usage of both source and target languages .
Approach: They propose a model that can model lyric translation and lyrics-melody alignment . they use an encoder-decoder framework that can translate lyrics and determine number of aligned notes .
Outcome: The proposed framework can translate lyrics and determine the number of aligned notes at each decoding step.
Multilingual Synopses of Movie Narratives: A Dataset for Vision-Language Story Understanding (2024.findings-emnlp)

Copied to clipboard

Challenge: Story video-text alignment is a core task in computational story understanding, but its progress has been held back by the scarcity of manually annotated video- text correspondences and the heavy concentration on English narrations of Hollywood movies.
Approach: They construct a multilingual video story dataset with 13,166 movie summary videos from 7 languages and manual annotations of fine-grained video-text correspondences.
Outcome: The proposed approach outperforms the SOTA methods on clip accuracy and Sentence IoU scores.
Sing it, Narrate it: Quality Musical Lyrics Translation (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing song translation approaches prioritize singability constraints at the expense of translation quality, which is crucial for musicals.
Approach: They propose to automatically translate musical lyrics from English to Chinese to ensure high translation quality while adhering to singability requirements such as length and rhyme.
Outcome: The proposed method improves both singability and translation quality over baseline methods and validates its effectiveness.
Towards Singable Lyrics Translation Using Large Language Models (2026.eacl-srw)

Copied to clipboard

Challenge: Existing studies on lyrics translation have relied on fine-tuning open-source language models.
Approach: They examine a multilingual lyrics translation dataset and apply prompting methods to large language models to evaluate singability.
Outcome: The proposed methods improve singability and naturalness, compared to naive translation, the authors show . human evaluations using songs created from translated lyrics show that complex prompting strategies improve singable naturalness .
Movie101v2: Improved Movie Narration Benchmark (2025.acl-long)

Copied to clipboard

Challenge: Automatic movie narration aims to generate video-aligned plot descriptions to assist visually impaired audiences.
Approach: They propose to break down the ultimate goal of automatic movie narration into three stages . they propose a large-scale, bilingual dataset with enhanced data quality .
Outcome: The proposed dataset breaks down the goal of automatic movie narration into three stages . achieving applicable movie narration is a fascinating goal that requires significant research .
K-pop Lyric Translation: Dataset, Analysis, and Neural-Modelling (2024.lrec-main)

Copied to clipboard

Challenge: lyric translation studies have focused on Western genres and languages, with no previous study centering on K-pop despite its popularity.
Approach: They propose a singable lyric translation dataset that aligns Korean and English lyrics line-by-line and section-by section.
Outcome: The proposed dataset reveals unique characteristics of K-pop lyric translation, distinguishing it from other extensively studied genres, and constructs a neural lyrical translation model.
SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition (2025.acl-long)

Copied to clipboard

Challenge: Creating lyrics and melodies in symbolic format requires expert knowledge of melody and an advanced understanding of lyrics.
Approach: They introduce SongComposer, a music-specialized large language model that can create symbolic lyrics and melodies following instructions.
Outcome: The proposed model outperforms existing models in symbolic song composition tasks.
REFFLY: Melody-Constrained Lyrics Editing Model (2025.naacl-long)

Copied to clipboard

Challenge: Automatic melody-to-lyric (M2L) generation aims to create lyrics that align with a given melody.
Approach: They propose a framework for automatic melody-to-lyric generation that allows for a more flexible approach to creating lyrics from plain text.
Outcome: The proposed framework outperforms baselines Lyra and GPT-4 in musicality and text quality.
Automatic Song Translation for Tonal Languages (2022.findings-acl)

Copied to clipboard

Challenge: Existing automatic song translation systems for tonal languages do not match the number of notes and beat the original rhythm of the song.
Approach: They propose three criteria for effective AST: preserving meaning, singability and intelligibility.
Outcome: The proposed system balances semantics and singability with human evaluations.
A Melody-Conditioned Lyrics Language Model (N18-1)

Copied to clipboard

Challenge: Existing models for lyrics generation are insufficient to capture relationship between lyrics and melody.
Approach: They propose a data-driven language model that generates entire lyrics for a given melody.
Outcome: The proposed model generates fluent lyrics while maintaining compatibility between lyrics and melodies.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations