In-the-wild Audio Spatialization with Flexible Text-guided Localization (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for mapping monaural audio to binaural signals lack flexibility and interactive control needed in complex multi-object user-interactive environments. |
| Approach: | They propose a text-guided audio spatialization framework that utilizes diverse text prompts to evaluate binaural audio models. |
| Outcome: | The proposed framework learns binaural differences guided by 3D spatial location and relative position prompts, enhanced with flipped-channel audio. |
Similar Papers
ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling (2026.acl-long)
Copied to clipboard
| Challenge: | Recent efforts on text-to-audio generation are exploring fine-grained controllability . however, their performance at scale is limited due to data scarcity . |
| Approach: | They propose a multi-task learning problem for high-controllability text-to-audio generation . they propose scalable diffusion transformers that augment condition information in sequence . |
| Outcome: | The proposed method outperforms existing methods on objective and subjective evaluations. |
On the Semantic Latent Space of Diffusion-Based Text-To-Speech Models (2024.acl-short)
Copied to clipboard
| Challenge: | Denoising Diffusion Models (DDMs) are a powerful generative tool for text-to-speech (TTS) but their semantic capabilities are unknown and control of synthesized speech’s vocal properties remains a challenge. |
| Approach: | They explore the latent space of frozen TTS models composed of latent bottleneck activations of the DDM’s denoiser and propose methods for finding semantic directions within it. |
| Outcome: | The proposed methods enable off-the-shelf audio editing without any training, architectural changes or data requirements. |
Beyond Transcription: Unified Audio Schema for Perception-Aware AudioLLMs (2026.findings-acl)
Copied to clipboard
Linhao Zhang, Yuhan Song, Aiwei Liu, Chuhan Wu, Sijun Zhang, Wei Jia, Yuan Liu, Houfeng Wang, Zhou Xiao
| Challenge: | Recent Audio Large Language Models (AudioLLMs) excel at reasoning tasks, but struggle at elementary auditory perception. |
| Approach: | They propose a framework that organizes audio information into three explicit components in a unified JSON format. |
| Outcome: | The proposed framework boosts fine-grained perception by 10.9% on MMSU over state-of-the-art models while preserving robust reasoning capabilities. |
ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment (2026.acl-long)
Copied to clipboard
| Challenge: | ImmersiveTTS model synthesizes intelligible speech and environmental audio from natural language descriptions. |
| Approach: | They propose an environment-aware text-to-speech model that integrates natural speech with environmental audio . the model explicitly models cross-modal interactions through a dual-stream stage . |
| Outcome: | Experimental results show that ImmersiveTTS achieves higher naturalness, intelligibility, and audio fidelity than existing approaches. |
Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluations of audio large language models focus on single audio inputs, but real-world applications often require processing multiple audio streams simultaneously. |
| Approach: | They propose a multi-audio evaluation benchmark that combines 20 audio inputs from 11 audio tasks to capture audio context. |
| Outcome: | The proposed model outperforms baseline models and achieves high data efficiency without human annotations. |
The Sonar Moment: An Audio Geo-Localization Benchmark for Audio-Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | AGL1K is the first audio geo-localization benchmark for audio language models (ALMs) it is based on a crowd-sourced platform and is available in 72 countries and territories. |
| Approach: | They propose a benchmark for audio geo-localization that quantifies the informativeness of each recording and a metric that quantizes the information of each audio clip. |
| Outcome: | The proposed benchmarks cover 72 countries and territories and can be used to improve audio geo-localization. |
Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for audio captioning lack fine-grained detail and contextual accuracy due to limited unimodal or superficial information. |
| Approach: | They propose a two-stage automated pipeline that uses pretrained models to extract contextual cues from video . a large language model synthesizes these inputs to generate detailed and context-aware captions . |
| Outcome: | The proposed method is scalable and generates detailed and context-aware captions on large-scale audio datasets. |
AudioCaps: Generating Captions for Audios in The Wild (N19-1)
Copied to clipboard
| Challenge: | a dataset of 46K audio clips with human-written text pairs is used to generate captions for audio . the task of translating a multimedia input source into natural language has been extensively studied over the past few years . |
| Approach: | They propose a top-down multi-scale encoder and aligned semantic attention for audio captioning. |
| Outcome: | The proposed captions are faithful to audio inputs and better than existing models. |
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback (2025.acl-long)
Copied to clipboard
Zehan Wang, Ke Lei, Chen Zhu, Jiawei Huang, Sashuai Zhou, Luping Liu, Xize Cheng, Shengpeng Ji, Zhenhui Ye, Tao Jin, Zhou Zhao
| Challenge: | Text-to-audio (T2A) models still struggle to satisfy human preferences for prompt-following and acoustic quality when generating complex multi-event audio. |
| Approach: | They propose to use AI feedback learning to enhance basic capabilities of text-to-audio models . they use a large audio preference dataset to evaluate the model's capabilities . |
| Outcome: | The proposed model improves in simple and complex scenarios with AI feedback learning. |
TAVT: Towards Transferable Audio-Visual Text Generation (2023.acl-long)
Copied to clipboard
| Challenge: | Existing transfer learning techniques focus on uni-modal analysis and lack consideration of multi-modal content and cross-modal relation. |
| Approach: | They propose a transferable audio-visual text generation framework that incorporates two components: Audio-Visual Meta-Mapper and Dual Counterfactual Contrastive Learning. |
| Outcome: | The proposed framework outperforms the state-of-the-art methods across multiple domains and modal settings. |