Challenge: We propose a novel large-scale audio-language model with advanced audio understanding and reasoning abilities.
Approach: They propose a general-purpose large audio-language model with advanced audio understanding and reasoning abilities that integrates an LLM with multiple types of audio representations.
Outcome: The proposed model outperforms existing models on audio understanding tasks by 1%-84%.

Similar Papers

Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large audio-language models (LALMs) have expanded their impact beyond natural language processing (NLP) to multimodal domains.
Approach: They propose a systematic taxonomy for LALM evaluations, categorizing them into four dimensions based on their objectives: (1) General Auditory Awareness and Processing, (2) Knowledge and Reasoning, (3) Dialogue-oriented Ability, and (4) Fairness, Safety, and Trustworthiness.
Outcome: The proposed taxonomy categorizes LALM evaluations into four dimensions based on their objectives and highlights challenges in this field.
AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs (2024.naacl-long)

Copied to clipboard

Challenge: a new model for speech processing and reasoning uses curated data instead of text.
Approach: They extend the instruction-tuned Llama-2 model with end-to-end speech processing and reasoning abilities without using any carefully curated paired data.
Outcome: The proposed model outperforms or outperfects existing models on synthesized and recorded speech QA tests.
Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in multimodal reasoning overlook the audio modality.
Approach: They propose a large-scale audio language model for deep reasoning that leverages a multitask audio dataset.
Outcome: The proposed model performs well across key benchmarks including MMAU-mini, AIR-Bench chat/foundation, and MELD.
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for audio-centric interaction have impeded advancements in this field . AIR-Bench evaluates LALMs' ability to understand audio signals and interact with humans .
Approach: They propose a benchmark to evaluate the ability of large audio-language models to understand audio signals . they use 19 tasks with approximately 19k single-choice questions to examine single-task ability .
Outcome: The proposed framework evaluates the ability of large audio-language models to understand audio signals and interact with humans in the textual format.
Enhancing Temporal Understanding in Audio Question Answering for Large Audio Language Models (2025.naacl-industry)

Copied to clipboard

Challenge: Recent literature focuses on constructing large audio language models (LALMs) but they are limited in temporal reasoning, which may hinder commercial applications .
Approach: They propose a data augmentation technique for generating reliable audio temporal questions and answers using an LLM.
Outcome: The proposed model performs well on public audio benchmark datasets and is optimized for edge applications.
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing benchmarks for testing audio modality of multimodal large language models focus on testing audio tasks in isolation.
Approach: They propose a new benchmark to assess multimodal large language models' ability to combine audio tasks.
Outcome: The proposed benchmarks show that multimodal models can solve problems that require reasoning over audio signals with satisfactory results.
PolyAudio: Advancing Multi-Audio Reasoning in Large Audio Language Models with Interleaved Multi-Audio Contexts (2026.findings-acl)

Copied to clipboard

Challenge: Large Audio Language Models have shown impressive performance on single-clip tasks . however, their ability to reason over interleaved multi-audio contexts remains limited .
Approach: They propose a LALM that targets multi-audio understanding via instruction tuning rather than massive-scale pre-training.
Outcome: The proposed model outperforms baseline models on multi-audio tasks while maintaining robustness.
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases (2026.eacl-long)

Copied to clipboard

Challenge: Existing models that measure audio comprehension beyond automatic speech recognition lack performance and latency.
Approach: They propose a benchmark suite that measures audio comprehension beyond automatic speech recognition . the benchmark suite includes a small human-recorded evaluation split per category .
Outcome: The proposed suite measures audio comprehension beyond speech recognition . it includes a small human-recorded evaluation split per category .
Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluations of audio large language models focus on single audio inputs, but real-world applications often require processing multiple audio streams simultaneously.
Approach: They propose a multi-audio evaluation benchmark that combines 20 audio inputs from 11 audio tasks to capture audio context.
Outcome: The proposed model outperforms baseline models and achieves high data efficiency without human annotations.
SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent large language models have demonstrated impressive reasoning abilities, but their extension to the audio modality remains underexplored.
Approach: They propose a rule-based reinforcement learning algorithm to equip LALMs with robust reasoning capabilities.
Outcome: The proposed algorithm improves on the SoundMind benchmark.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations