QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions (2025.acl-long)
Copied to clipboard
Siyin Wang, Wenyi Yu, Xianzhao Chen, Xiaohai Tian, Jun Zhang, Lu Lu, Yu Tsao, Junichi Yamagishi, Yuxuan Wang, Chao Zhang
| Challenge: | Existing datasets lack comprehensive annotations for speech quality assessment . existing methods lack detailed annotations, resulting in inaccurate evaluations. |
| Approach: | They propose a low-level speech quality assessment dataset incorporating natural language descriptions and a Benchmark to evaluate low- level speech understanding capabilities of auditory large language models. |
| Outcome: | The proposed model can be used to evaluate the low-level speech understanding capabilities of auditory large language models. |
Similar Papers
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation (2026.acl-long)
Copied to clipboard
Hui Wang, Jinghua Zhao, Yifan Yang, Shujie Liu, Junyang Chen, Yanzhe Zhang, Shiwan Zhao, Jinyu Li, Jiaming Zhou, Haoqin Sun, Yan Lu, Yong Qin
| Challenge: | Existing methods for evaluating the perceptual quality of synthetic speech are limited due to the complexity of perceptual quality factors and the diversity of speech generation tasks. |
| Approach: | They propose a new paradigm for enabling large language models to conduct structured speech quality evaluation using a large-scale dataset. |
| Outcome: | The proposed model performs well across tasks and languages. |
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing models that measure audio comprehension beyond automatic speech recognition lack performance and latency. |
| Approach: | They propose a benchmark suite that measures audio comprehension beyond automatic speech recognition . the benchmark suite includes a small human-recorded evaluation split per category . |
| Outcome: | The proposed suite measures audio comprehension beyond speech recognition . it includes a small human-recorded evaluation split per category . |
WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem (2026.findings-acl)
Copied to clipboard
Chengyou Wang, Mingchen Shao, Jingbin Hu, Zeyu Zhu, Hongfei Xue, Bingshen Mu, Xin Xu, Xingyi Duan, Binbin Zhang, Zhu Pengcheng, Chuang Ding, Xiaojun Zhang, Hui Bu, Lei Xie
| Challenge: | despite its linguistic significance, the Wu dialect of Chinese has long been hindered by the lack of large-scale speech data, standardized evaluation benchmarks, and publicly available models. |
| Approach: | They propose to use WenetSpeech-Wu as a large-scale, multi-dimensionally annotated open-source speech corpus for the Wu dialect of Chinese. |
| Outcome: | The proposed dataset includes 8,000 hours of speech data and strong open-source models . the proposed dataset is competitive and empirically validated . |
VoiceBench: Benchmarking LLM-Based Voice Assistants (2026.tacl-1)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have enabled real-time speech interactions through LLMs. |
| Approach: | They propose a benchmark specifically designed to assess LLM-based voice assistants. |
| Outcome: | The proposed benchmark measures the performance of LLM-based voice assistants across eight tasks. |
QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing benchmarks often lack domain coverage and provide limited insights into the working context of Chinese LLMs. |
| Approach: | They propose a multi-domain Chinese QA benchmark dedicated to localized assessment of Chinese LLMs. |
| Outcome: | The Qwen2.5 model outperforms the more advanced GPT-4o model in the Chinese market . the dataset includes over 17,000 questions across six vertical domains . |
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation datasets lack cross-lingual alignment, leaving assessments of multilingual capabilities fragmented in both language and skill coverage. |
| Approach: | They propose to use multilingual consistency as a complementary metric to assess performance bottlenecks and guide model improvement. |
| Outcome: | The proposed model lacks cross-lingual alignment and language coverage gaps between state-of-the-art models. |
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing question-answering benchmarks fail to evaluate SLMs’ knowledge understanding due to their inability to support end-to-end speech evaluation and account for varied input audio conditions. |
| Approach: | They propose a new question-answering benchmark that assesses SLMs’ knowledge understanding through pure speech interactions. |
| Outcome: | The proposed benchmark maintains speech format for both inputs and outputs, evaluates model robustness across diverse input audio conditions, and pioneers the assessment of complex tasks like mathematical reasoning in spoken format. |
AudioBench: A Universal Benchmark for Audio Large Language Models (2025.naacl-long)
Copied to clipboard
Bin Wang, Xunlong Zou, Geyu Lin, Shuo Sun, Zhuohan Liu, Wenyu Zhang, Zhengyuan Liu, AiTi Aw, Nancy F. Chen
| Challenge: | Existing evaluation regimes for audio large language models do not cover the breadth of their possible use cases. |
| Approach: | They propose to use AudioBench to evaluate audio large language models . they found that no single model excels consistently across all tasks . |
| Outcome: | The proposed evaluation targets speech understanding, audio scene understanding, and voice understanding (paralinguistic) . no single model excels consistently across all tasks, the paper found . |
CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs (2025.naacl-long)
Copied to clipboard
| Challenge: | Current evaluation methods do not capture complex attributes of counterspeech quality, such as contextual relevance, aggressiveness, or argumentative coherence. |
| Approach: | They propose to use a dataset and framework to evaluate counterspeech quality across four dimensions: contextual relevance, aggressiveness, argument-coherence, and suitability. |
| Outcome: | The proposed method outperforms ROUGE, METEOR, and BertScore in correlating with human judgement, indicating a significant improvement in automated counterspeech evaluation. |
From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment (2026.findings-acl)
Copied to clipboard
Ivo Bueno, Babette Bühler, Philipp Stark, Tim Fütterer, Ulrich Trautwein, Dorottya Demszky, Heather Hill, Enkelejda Kasneci
| Challenge: | a framework for sentence-level interpretability of rubric-based scoring is proposed . aaron e. smith: automated scoring models provide little insight into why scores are produced . |
| Approach: | They propose a framework for sentence-level interpretability of rubric-based scoring that combines Shapley-value attributions with rationales generated by large language models. |
| Outcome: | The proposed framework compares fine-tuned pretrained language models with large language models . it shows that fine- tuned models outperform LLMs in prediction accuracy but exhibit label compression toward mid-scale scores . |