Afrispeech Semantics: Evaluating Audio–Semantic Reasoning in Spoken Language Models Across Domains and Accents (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent multimodal models are trained on large collections of audio-text pairs using contrastive learning or nexttoken prediction objectives. |
| Approach: | They evaluate audio language models across five semantic and paralinguistic reasoning tasks: entailment, consistency, plausibility, accent drift, and accent restraint. |
| Outcome: | The evaluations assess models across five tasks including entailment, consistency, plausibility, accent drift, and accent restraint. |
Similar Papers
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models (2026.eacl-long)
Copied to clipboard
Iwona Christop, Mateusz Czyżnikiewicz, Paweł Skórzewski, Łukasz Bondaruk, Jakub Kubiak, Marcin Lewandowski, Marek Kubis
| Challenge: | Existing benchmarks for testing audio modality of multimodal large language models focus on testing audio tasks in isolation. |
| Approach: | They propose a new benchmark to assess multimodal large language models' ability to combine audio tasks. |
| Outcome: | The proposed benchmarks show that multimodal models can solve problems that require reasoning over audio signals with satisfactory results. |
Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in large audio-language models (LALMs) have expanded their impact beyond natural language processing (NLP) to multimodal domains. |
| Approach: | They propose a systematic taxonomy for LALM evaluations, categorizing them into four dimensions based on their objectives: (1) General Auditory Awareness and Processing, (2) Knowledge and Reasoning, (3) Dialogue-oriented Ability, and (4) Fairness, Safety, and Trustworthiness. |
| Outcome: | The proposed taxonomy categorizes LALM evaluations into four dimensions based on their objectives and highlights challenges in this field. |
Hearing Between the Lines: Unlocking the Reasoning Power of LLMs for Speech Evaluation (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large Language Model (LLM) judges are limited to textual content, resulting in expensive and opaque evaluation methods. |
| Approach: | They propose a framework that enables large language model judges to reason over audio cues . they introduce a human chain-of-thought annotation protocol to improve judge diagnostic capability . |
| Outcome: | The proposed framework achieves higher agreement with human raters than ALMs and transcript-only LLM judges while being significantly more cost-effective. |
Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance (2026.eacl-long)
Copied to clipboard
| Challenge: | LISTEN is a controlled benchmark to disentangle lexical reliance from acoustic sensitivity in emotion understanding. |
| Approach: | They propose a benchmark to disentangle lexical reliance from acoustic sensitivity in emotion understanding. |
| Outcome: | LISTEN shows that current LALMs largely "transcribe" rather than "listen" authors note that models underutilize acoustic cues while relying on lexical semantics . |
Speech language models lack important brain-relevant semantics (2024.acl-long)
Copied to clipboard
| Challenge: | Recent work shows that text-based language models predict both text- and speech-evoked brain activity. |
| Approach: | They remove low-level stimulus features from language models to assess their impact on alignment with fMRI brain recordings during reading and listening. |
| Outcome: | The proposed model removes low-level features from fMRI brain recordings to assess their impact on alignment with fmr recordings. |
Evaluating Open-Source ASR Systems: Performance Across Diverse Audio Conditions and Error Correction Methods (2025.coling-main)
Copied to clipboard
| Challenge: | Automated speech recognition (ASR) systems are able to transcribe spontaneous human conversations with high accuracy. |
| Approach: | They evaluate the accuracy of open source automatic speech recognition systems across conversational speech datasets and explore the potential of ASR ensembling and post-ASR correction methods to improve transcription accuracy. |
| Outcome: | The proposed methods highlight the need for robust error correction techniques and address demographic biases to enhance ASR performance and inclusivity. |
Explaining Speech Classification Models via Word-Level Audio Segments and Paralinguistic Features (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing explanations for speech classification models are difficult to interpret and make mistakes. |
| Approach: | They propose to explain speech classification models by using word-level and paralinguistic attributes to measure the impact of each audio segment aligned with a word on the outcome. |
| Outcome: | The proposed explanations correctly represent the model’s inner workings and are plausible to humans. |
Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations (2026.findings-eacl)
Copied to clipboard
| Challenge: | Recent advances extend language understanding beyond text to speech, enabling unified reasoning across modalities. |
| Approach: | They construct and release a speech-augmented benchmark based on Global MMLU Lite and a data set spanning English, Chinese, and Korean. |
| Outcome: | The proposed model is robust to demographic factors but sensitive to language and option order, suggesting that speech can amplify structural biases. |
SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues? (2026.acl-long)
Copied to clipboard
| Challenge: | Large Audio-Language Models (LALMs) are a popular approach for evaluating speech quality, yet their ability to assess speaker consistency across multi-turn dialogues remains unexplored. |
| Approach: | They construct 1,818 human-verified evaluation instances across four datasets spanning synthetic and real speech, with controlled acoustic difficulty. |
| Outcome: | The proposed model performs better in comparing and ranking acoustic variants, demonstrating inherent acustic discrimination capabilities. |
Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in multimodal reasoning overlook the audio modality. |
| Approach: | They propose a large-scale audio language model for deep reasoning that leverages a multitask audio dataset. |
| Outcome: | The proposed model performs well across key benchmarks including MMAU-mini, AIR-Bench chat/foundation, and MELD. |