RoboVox: A Single/Multi-channel Far-field Speaker Recognition Benchmark for a Mobile Robot (2024.lrec-main)
Copied to clipboard
Mohammad Mohammadamini, Driss Matrouf, Michael Rouvier, Jean-Francois Bonastre, Romain Serizel, Theophile Gonos
| Challenge: | In this paper, we introduce a new far-field speaker recognition benchmark called RoboVox. |
| Approach: | They introduce a new far-field speaker recognition benchmark called RoboVox which measures the far-feet of a French corpus recorded by a mobile robot. |
| Outcome: | The proposed benchmarks show a significant decline in far-field speaker recognition and urge the community to further research in this domain. |
Similar Papers
Far-Field Speaker Recognition Benchmark Derived From The DiPCo Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Using a publicly-available corpus, we propose a far-field speaker verification benchmark. |
| Approach: | They propose a far-field speaker verification benchmark derived from the publicly available DiPCo corpus. |
| Outcome: | The proposed tasks are very challenging and hope to inspire the speech community to develop new methods and systems for this challenging domain. |
ELITR-Bench: A Meeting Assistant Benchmark for Long-Context Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Existing benchmarks for long-context LLMs focus on generic tasks that are not necessarily aligned with real-world applications. |
| Approach: | They propose to augment existing ELITR corpus by adding 271 manually crafted questions with their ground-truth answers and noisy versions of meeting transcripts altered to target different Word Error Rate levels. |
| Outcome: | The proposed benchmark augments the existing ELITR corpus by adding 271 manually crafted questions with ground-truth answers, as well as noisy versions of meeting transcripts altered to target different Word Error Rate levels. |
AudioBench: A Universal Benchmark for Audio Large Language Models (2025.naacl-long)
Copied to clipboard
Bin Wang, Xunlong Zou, Geyu Lin, Shuo Sun, Zhuohan Liu, Wenyu Zhang, Zhengyuan Liu, AiTi Aw, Nancy F. Chen
| Challenge: | Existing evaluation regimes for audio large language models do not cover the breadth of their possible use cases. |
| Approach: | They propose to use AudioBench to evaluate audio large language models . they found that no single model excels consistently across all tasks . |
| Outcome: | The proposed evaluation targets speech understanding, audio scene understanding, and voice understanding (paralinguistic) . no single model excels consistently across all tasks, the paper found . |
VoiceBench: Benchmarking LLM-Based Voice Assistants (2026.tacl-1)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have enabled real-time speech interactions through LLMs. |
| Approach: | They propose a benchmark specifically designed to assess LLM-based voice assistants. |
| Outcome: | The proposed benchmark measures the performance of LLM-based voice assistants across eight tasks. |
VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents (2026.findings-acl)
Copied to clipboard
Jiliang Hu, Wenfu Wang, Zuchao Li, Chenxing Li, Yiyang Zhao, Hanzhao Li, Liqiang Zhang, Meng Yu, Dong Yu
| Challenge: | despite advances in multimodal conversational systems, current benchmarks lack comprehensive evaluation across key dimensions. |
| Approach: | They propose a Chinese benchmark built exclusively on real human speech to fill this gap . they assess LALMs across three complementary axes: instruction following, knowledge understanding, robustness . |
| Outcome: | VCB Bench assesses LALMs across three complementary axes: instruction following, knowledge understanding, and robustness . VCBM Bench provides reproducible and fine-grained framework for Chinese voice chat bots . results show significant performance disparities and offer tangible insights for future improvements . |
Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluations of audio large language models focus on single audio inputs, but real-world applications often require processing multiple audio streams simultaneously. |
| Approach: | They propose a multi-audio evaluation benchmark that combines 20 audio inputs from 11 audio tasks to capture audio context. |
| Outcome: | The proposed model outperforms baseline models and achieves high data efficiency without human annotations. |
MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions (2025.emnlp-main)
Copied to clipboard
Ramaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi, Sonal Kumar, Sreyan Ghosh, Dinesh Manocha
| Challenge: | omni models lack spoken dialogues, which is essential for assessing conversational and auditory capabilities of voice assistants. |
| Approach: | They propose a benchmark to evaluate the ability of voice assistants to integrate paralinguistic speech features into their models. |
| Outcome: | The multivox voice assistant benchmark evaluates the ability of models to integrate spoken and visual cues including paralinguistic speech features for truly multimodal understanding. |
Nkululeko: A Tool For Rapid Speaker Characteristics Detection (2022.lrec-1)
Copied to clipboard
| Challenge: | Nkululeko is a software tool that lets users perform semi-supervised machine learning experiments in the speaker characteristics domain. |
| Approach: | They propose a software tool called Nkululeko that lets users perform semi-supervised machine learning experiments in the speaker characteristics domain. |
| Outcome: | The proposed tool is based on audformat, a speech database metadata description . it supports best practise and fast setup of experiments without programming skills . |
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing models that measure audio comprehension beyond automatic speech recognition lack performance and latency. |
| Approach: | They propose a benchmark suite that measures audio comprehension beyond automatic speech recognition . the benchmark suite includes a small human-recorded evaluation split per category . |
| Outcome: | The proposed suite measures audio comprehension beyond speech recognition . it includes a small human-recorded evaluation split per category . |
XMAD-Bench: Cross-Domain Multilingual Audio Deepfake Benchmark (2026.findings-eacl)
Copied to clipboard
Ioan-Paul Ciobanu, Andrei-Iulian Hîji, Nicolae Catalin Ristea, Paul Irofti, Cristian Rusu, Radu Tudor Ionescu
| Challenge: | Recent advances in audio generation led to an increasing number of deepfakes . however, these methods are typically tested in an in-domain setup . |
| Approach: | They propose a large-scale cross-domain audio deepfake benchmark comprising 668.8 hours of real and deepfak speech. |
| Outcome: | The proposed benchmark compares audio deepfake detectors with existing methods in the wild . the results show that the proposed methods perform better in different languages than existing methods . |