Challenge: Existing evaluations of audio large language models focus on single audio inputs, but real-world applications often require processing multiple audio streams simultaneously.
Approach: They propose a multi-audio evaluation benchmark that combines 20 audio inputs from 11 audio tasks to capture audio context.
Outcome: The proposed model outperforms baseline models and achieves high data efficiency without human annotations.

Similar Papers

Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large audio-language models (LALMs) have expanded their impact beyond natural language processing (NLP) to multimodal domains.
Approach: They propose a systematic taxonomy for LALM evaluations, categorizing them into four dimensions based on their objectives: (1) General Auditory Awareness and Processing, (2) Knowledge and Reasoning, (3) Dialogue-oriented Ability, and (4) Fairness, Safety, and Trustworthiness.
Outcome: The proposed taxonomy categorizes LALM evaluations into four dimensions based on their objectives and highlights challenges in this field.
When Large Language Models Meet Speech: A Survey on Integration Approaches (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have spurred interest in expanding their application beyond text-based tasks.
Approach: They propose to categorize the integration of speech with LLMs into three main approaches . they demonstrate how these methods are applied across various speech-related applications .
Outcome: The proposed methods are applied across speech-related applications and highlight the challenges in this field to offer inspiration for future research.
PolyAudio: Advancing Multi-Audio Reasoning in Large Audio Language Models with Interleaved Multi-Audio Contexts (2026.findings-acl)

Copied to clipboard

Challenge: Large Audio Language Models have shown impressive performance on single-clip tasks . however, their ability to reason over interleaved multi-audio contexts remains limited .
Approach: They propose a LALM that targets multi-audio understanding via instruction tuning rather than massive-scale pre-training.
Outcome: The proposed model outperforms baseline models on multi-audio tasks while maintaining robustness.
Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask Learners (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have been used for general-purpose interfaces across multiple tasks and languages.
Approach: They propose to use large language models as a general-purpose interface across multiple tasks and languages.
Outcome: The proposed model performs better on 200K hours of 6-language data for voice generation applications.
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases (2026.eacl-long)

Copied to clipboard

Challenge: Existing models that measure audio comprehension beyond automatic speech recognition lack performance and latency.
Approach: They propose a benchmark suite that measures audio comprehension beyond automatic speech recognition . the benchmark suite includes a small human-recorded evaluation split per category .
Outcome: The proposed suite measures audio comprehension beyond speech recognition . it includes a small human-recorded evaluation split per category .
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing benchmarks for testing audio modality of multimodal large language models focus on testing audio tasks in isolation.
Approach: They propose a new benchmark to assess multimodal large language models' ability to combine audio tasks.
Outcome: The proposed benchmarks show that multimodal models can solve problems that require reasoning over audio signals with satisfactory results.
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for audio-centric interaction have impeded advancements in this field . AIR-Bench evaluates LALMs' ability to understand audio signals and interact with humans .
Approach: They propose a benchmark to evaluate the ability of large audio-language models to understand audio signals . they use 19 tasks with approximately 19k single-choice questions to examine single-task ability .
Outcome: The proposed framework evaluates the ability of large audio-language models to understand audio signals and interact with humans in the textual format.
LLMs for Low Resource Languages in Multilingual, Multimodal and Dialectal Settings (2024.eacl-tutorials)

Copied to clipboard

Challenge: Recent advances in AI can be attributed to the remarkable performance of Large Language Models (LLMs) success of LLMs depends on specific training techniques, such as instruction tuning and prompting .
Approach: They explore the capabilities of Large Language Models (LLMs) in various tasks and languages . they also examine their performance, fine-tuning, instructions tuning, and close vs. open models .
Outcome: The proposed model can be used for speech and multimodal tasks across modalities, languages, and dialects.
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) provide a data-centric solution to alleviate limitations of real-world data with synthetic data generation.
Approach: They propose a generic workflow for LLM-driven synthetic data generation.
Outcome: The proposed workflows highlight gaps in existing research and outline avenues for future studies.
Beyond Transcription: Unified Audio Schema for Perception-Aware AudioLLMs (2026.findings-acl)

Copied to clipboard

Challenge: Recent Audio Large Language Models (AudioLLMs) excel at reasoning tasks, but struggle at elementary auditory perception.
Approach: They propose a framework that organizes audio information into three explicit components in a unified JSON format.
Outcome: The proposed framework boosts fine-grained perception by 10.9% on MMSU over state-of-the-art models while preserving robust reasoning capabilities.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations