Challenge: In this paper, we introduce a new far-field speaker recognition benchmark called RoboVox.
Approach: They introduce a new far-field speaker recognition benchmark called RoboVox which measures the far-feet of a French corpus recorded by a mobile robot.
Outcome: The proposed benchmarks show a significant decline in far-field speaker recognition and urge the community to further research in this domain.

Similar Papers

Far-Field Speaker Recognition Benchmark Derived From The DiPCo Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Using a publicly-available corpus, we propose a far-field speaker verification benchmark.
Approach: They propose a far-field speaker verification benchmark derived from the publicly available DiPCo corpus.
Outcome: The proposed tasks are very challenging and hope to inspire the speech community to develop new methods and systems for this challenging domain.
ELITR-Bench: A Meeting Assistant Benchmark for Long-Context Language Models (2025.coling-main)

Copied to clipboard

Challenge: Existing benchmarks for long-context LLMs focus on generic tasks that are not necessarily aligned with real-world applications.
Approach: They propose to augment existing ELITR corpus by adding 271 manually crafted questions with their ground-truth answers and noisy versions of meeting transcripts altered to target different Word Error Rate levels.
Outcome: The proposed benchmark augments the existing ELITR corpus by adding 271 manually crafted questions with ground-truth answers, as well as noisy versions of meeting transcripts altered to target different Word Error Rate levels.
AudioBench: A Universal Benchmark for Audio Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Existing evaluation regimes for audio large language models do not cover the breadth of their possible use cases.
Approach: They propose to use AudioBench to evaluate audio large language models . they found that no single model excels consistently across all tasks .
Outcome: The proposed evaluation targets speech understanding, audio scene understanding, and voice understanding (paralinguistic) . no single model excels consistently across all tasks, the paper found .
VoiceBench: Benchmarking LLM-Based Voice Assistants (2026.tacl-1)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have enabled real-time speech interactions through LLMs.
Approach: They propose a benchmark specifically designed to assess LLM-based voice assistants.
Outcome: The proposed benchmark measures the performance of LLM-based voice assistants across eight tasks.
VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents (2026.findings-acl)

Copied to clipboard

Challenge: despite advances in multimodal conversational systems, current benchmarks lack comprehensive evaluation across key dimensions.
Approach: They propose a Chinese benchmark built exclusively on real human speech to fill this gap . they assess LALMs across three complementary axes: instruction following, knowledge understanding, robustness .
Outcome: VCB Bench assesses LALMs across three complementary axes: instruction following, knowledge understanding, and robustness . VCBM Bench provides reproducible and fine-grained framework for Chinese voice chat bots . results show significant performance disparities and offer tangible insights for future improvements .
Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluations of audio large language models focus on single audio inputs, but real-world applications often require processing multiple audio streams simultaneously.
Approach: They propose a multi-audio evaluation benchmark that combines 20 audio inputs from 11 audio tasks to capture audio context.
Outcome: The proposed model outperforms baseline models and achieves high data efficiency without human annotations.
MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions (2025.emnlp-main)

Copied to clipboard

Challenge: omni models lack spoken dialogues, which is essential for assessing conversational and auditory capabilities of voice assistants.
Approach: They propose a benchmark to evaluate the ability of voice assistants to integrate paralinguistic speech features into their models.
Outcome: The multivox voice assistant benchmark evaluates the ability of models to integrate spoken and visual cues including paralinguistic speech features for truly multimodal understanding.
Nkululeko: A Tool For Rapid Speaker Characteristics Detection (2022.lrec-1)

Copied to clipboard

Challenge: Nkululeko is a software tool that lets users perform semi-supervised machine learning experiments in the speaker characteristics domain.
Approach: They propose a software tool called Nkululeko that lets users perform semi-supervised machine learning experiments in the speaker characteristics domain.
Outcome: The proposed tool is based on audformat, a speech database metadata description . it supports best practise and fast setup of experiments without programming skills .
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases (2026.eacl-long)

Copied to clipboard

Challenge: Existing models that measure audio comprehension beyond automatic speech recognition lack performance and latency.
Approach: They propose a benchmark suite that measures audio comprehension beyond automatic speech recognition . the benchmark suite includes a small human-recorded evaluation split per category .
Outcome: The proposed suite measures audio comprehension beyond speech recognition . it includes a small human-recorded evaluation split per category .
XMAD-Bench: Cross-Domain Multilingual Audio Deepfake Benchmark (2026.findings-eacl)

Copied to clipboard

Challenge: Recent advances in audio generation led to an increasing number of deepfakes . however, these methods are typically tested in an in-domain setup .
Approach: They propose a large-scale cross-domain audio deepfake benchmark comprising 668.8 hours of real and deepfak speech.
Outcome: The proposed benchmark compares audio deepfake detectors with existing methods in the wild . the results show that the proposed methods perform better in different languages than existing methods .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations