Challenge: Existing methods for mapping monaural audio to binaural signals lack flexibility and interactive control needed in complex multi-object user-interactive environments.
Approach: They propose a text-guided audio spatialization framework that utilizes diverse text prompts to evaluate binaural audio models.
Outcome: The proposed framework learns binaural differences guided by 3D spatial location and relative position prompts, enhanced with flipped-channel audio.

Similar Papers

ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling (2026.acl-long)

Copied to clipboard

Challenge: Recent efforts on text-to-audio generation are exploring fine-grained controllability . however, their performance at scale is limited due to data scarcity .
Approach: They propose a multi-task learning problem for high-controllability text-to-audio generation . they propose scalable diffusion transformers that augment condition information in sequence .
Outcome: The proposed method outperforms existing methods on objective and subjective evaluations.
On the Semantic Latent Space of Diffusion-Based Text-To-Speech Models (2024.acl-short)

Copied to clipboard

Challenge: Denoising Diffusion Models (DDMs) are a powerful generative tool for text-to-speech (TTS) but their semantic capabilities are unknown and control of synthesized speech’s vocal properties remains a challenge.
Approach: They explore the latent space of frozen TTS models composed of latent bottleneck activations of the DDM’s denoiser and propose methods for finding semantic directions within it.
Outcome: The proposed methods enable off-the-shelf audio editing without any training, architectural changes or data requirements.
Beyond Transcription: Unified Audio Schema for Perception-Aware AudioLLMs (2026.findings-acl)

Copied to clipboard

Challenge: Recent Audio Large Language Models (AudioLLMs) excel at reasoning tasks, but struggle at elementary auditory perception.
Approach: They propose a framework that organizes audio information into three explicit components in a unified JSON format.
Outcome: The proposed framework boosts fine-grained perception by 10.9% on MMSU over state-of-the-art models while preserving robust reasoning capabilities.
ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment (2026.acl-long)

Copied to clipboard

Challenge: ImmersiveTTS model synthesizes intelligible speech and environmental audio from natural language descriptions.
Approach: They propose an environment-aware text-to-speech model that integrates natural speech with environmental audio . the model explicitly models cross-modal interactions through a dual-stream stage .
Outcome: Experimental results show that ImmersiveTTS achieves higher naturalness, intelligibility, and audio fidelity than existing approaches.
Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluations of audio large language models focus on single audio inputs, but real-world applications often require processing multiple audio streams simultaneously.
Approach: They propose a multi-audio evaluation benchmark that combines 20 audio inputs from 11 audio tasks to capture audio context.
Outcome: The proposed model outperforms baseline models and achieves high data efficiency without human annotations.
The Sonar Moment: An Audio Geo-Localization Benchmark for Audio-Language Models (2026.findings-acl)

Copied to clipboard

Challenge: AGL1K is the first audio geo-localization benchmark for audio language models (ALMs) it is based on a crowd-sourced platform and is available in 72 countries and territories.
Approach: They propose a benchmark for audio geo-localization that quantifies the informativeness of each recording and a metric that quantizes the information of each audio clip.
Outcome: The proposed benchmarks cover 72 countries and territories and can be used to improve audio geo-localization.
Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for audio captioning lack fine-grained detail and contextual accuracy due to limited unimodal or superficial information.
Approach: They propose a two-stage automated pipeline that uses pretrained models to extract contextual cues from video . a large language model synthesizes these inputs to generate detailed and context-aware captions .
Outcome: The proposed method is scalable and generates detailed and context-aware captions on large-scale audio datasets.
AudioCaps: Generating Captions for Audios in The Wild (N19-1)

Copied to clipboard

Challenge: a dataset of 46K audio clips with human-written text pairs is used to generate captions for audio . the task of translating a multimedia input source into natural language has been extensively studied over the past few years .
Approach: They propose a top-down multi-scale encoder and aligned semantic attention for audio captioning.
Outcome: The proposed captions are faithful to audio inputs and better than existing models.
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback (2025.acl-long)

Copied to clipboard

Challenge: Text-to-audio (T2A) models still struggle to satisfy human preferences for prompt-following and acoustic quality when generating complex multi-event audio.
Approach: They propose to use AI feedback learning to enhance basic capabilities of text-to-audio models . they use a large audio preference dataset to evaluate the model's capabilities .
Outcome: The proposed model improves in simple and complex scenarios with AI feedback learning.
TAVT: Towards Transferable Audio-Visual Text Generation (2023.acl-long)

Copied to clipboard

Challenge: Existing transfer learning techniques focus on uni-modal analysis and lack consideration of multi-modal content and cross-modal relation.
Approach: They propose a transferable audio-visual text generation framework that incorporates two components: Audio-Visual Meta-Mapper and Dual Counterfactual Contrastive Learning.
Outcome: The proposed framework outperforms the state-of-the-art methods across multiple domains and modal settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations