MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks focus on evaluating single-round interactions, neglecting other critical aspects. |
| Approach: | They propose a benchmark to evaluate full-duplex speech language models in multi-round settings . they segment continuous full-dual dialogues into discrete turns for evaluation . |
| Outcome: | The proposed benchmark compared full-duplex speech language models with full-dual speech models . the results show that the models perform better in multi-round settings than standard models compared to benchmarks . |
Similar Papers
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation (2026.acl-long)
Copied to clipboard
Xiaoyuan Li, Keqin Bao, Yubo Ma, Moxin Li, Wenjie Wang, Rui Men, Yichang Zhang, Fuli Feng, Dayiheng Liu
| Challenge: | Recent advances in Large Language Models (LLMs) have shown promising results in complex reasoning tasks. |
| Approach: | They propose to use a multi-turn reasoning evaluation framework to cover multi-turn interactions with the environments of large language models. |
| Outcome: | The proposed framework covers diverse reasoning capabilities, fine-grained difficulty granularity, and necessitates multi-turn interactions with the environments. |
MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues (2024.acl-long)
Copied to clipboard
Ge Bai, Jie Liu, Xingyuan Bu, Yancheng He, Jiaheng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, Wanli Ouyang
| Challenge: | Large Language Models (LLMs) have greatly enhanced dialogue systems, but evaluation of their capabilities remains a challenge. |
| Approach: | They propose a model to evaluate the fine-grained abilities of Large Language Models in multi-turn dialogues. |
| Outcome: | The proposed model evaluates 21 popular chatbots based on MT-Bench-101 . it includes 3 overarching abilities and 13 distinct tasks within multi-turn dialogue scenarios. |
Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner (2026.acl-short)
Copied to clipboard
Guan-Ting Lin, Shih-Yun Shan Kuan, Jiatong Shi, Kai-Wei Chang, Siddhant Arora, Shinji Watanabe, Hung-yi Lee
| Challenge: | Full-duplex speech agents are often half-duplice, alternating turns between user and system. |
| Approach: | They propose a streaming framework that integrates with an examiner that enforces staged goals under two pacing setups. |
| Outcome: | The framework reports fluency, multi-turn instruction following, and task-specific competence. |
MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models (2024.emnlp-main)
Copied to clipboard
Wai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang, Liangyou Li, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong
| Challenge: | Existing evaluation frameworks focus on single-turn evaluations, overlooking the models’ capabilities in multi-turn interactions. |
| Approach: | They propose a benchmark to evaluate the multi-turn conversational abilities of large language models (LLMs) by analyzing human-LLM conversations and constructing multi-turned queries for each category using GPT-4. |
| Outcome: | The proposed model outperforms open-source models in multi-turn tasks while retaining and recalling historical information. |
Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models (2024.emnlp-main)
Copied to clipboard
Xinrong Zhang, Yingfa Chen, Shengding Hu, Xu Han, Zihang Xu, Yuanwei Xu, Weilin Zhao, Maosong Sun, Zhiyuan Liu
| Challenge: | Large language models (LLMs) are increasingly permeating daily lives and require real-time interactions that mirror human conversations. |
| Approach: | They propose to use time-division-multiplexing to process queries and responses pseudo-simultaneously. |
| Outcome: | The proposed model can listen to users while generating output and adjust to provide instant feedback. |
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing function-calling benchmarks focus on single-turn interactions but ignore complexity of real-world scenarios. |
| Approach: | They propose a framework that constructs practical function-calling datasets by synthesizing conversations through a tool graph that maintains dependencies across rounds. |
| Outcome: | The proposed framework synthesizes conversations through a tool graph that maintains dependencies across rounds and a multi-agent system with distinct personas to enhance dialogue naturalness. |
Sommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models (2026.acl-industry)
Copied to clipboard
| Challenge: | Existing high-quality conversational data is limited for full-duplex models . overlapping and backchanneling are a challenge for most systems . |
| Approach: | They propose a robust and scalable open-source data processing pipeline for full-duplex models. |
| Outcome: | The proposed pipeline can listen and speak simultaneously, supporting more fluid and human-like interaction. |
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues (2026.findings-acl)
Copied to clipboard
Yaning Pan, Qianqian Xie, Guohui Zhang, Zekun Moore Wang, Yongqian Wen, Yuanxing Zhang, Haoxuan Hu, Zhiyu Pan, Yibing Huang, Zhidong Gan, Yonghong Lin, An Ping, Shihao Li, Yanghai Wang, Tianhao Peng, Jiaheng Liu
| Challenge: | Existing evaluation benchmarks for Multimodal Large Language Models (MLLMs) focus on single-turn question answering, overlooking the complexity of multi-turn dialogues in real-world scenarios. |
| Approach: | They propose a video understanding benchmark for MLLMs in multi-turn dialogues that assesses six core competencies that focus on perceptivity and interactivity. |
| Outcome: | The MT-Video-Bench evaluates 1,000 multi-turn dialogues from diverse domains and reveals significant performance discrepancies and limitations in handling multi-turned video dialogues. |
F-Actor: Controllable Conversational Behavior in Full-Duplex Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Current spoken conversational systems lack customization capabilities, limiting their naturalness and usability. |
| Approach: | They propose an instruction-following full-duplex conversational speech model that can be trained efficiently under typical academic resource constraints. |
| Outcome: | The proposed model requires just 2,000 hours of data to be trained under typical academic resource constraints. |
TurnBench-MS: A Benchmark for Evaluating Multi-Turn, Multi-Step Reasoning in Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing benchmarks focus on single-turn or single-step tasks, failing to capture iterative reasoning in real-world settings. |
| Approach: | They propose a benchmark that evaluates multi-turn, multi-step reasoning through an interactive code-breaking task inspired by the "Turing Machine Board Game" the best model achieves 84% accuracy in Classic mode, but performance drops to 18% in Nightmare mode. |
| Outcome: | The new benchmark evaluates multi-turn, multi-step reasoning through an interactive code-breaking task inspired by the "Turing Machine Board Game" the best model achieves 84% accuracy in Classic mode, but performance drops to 18% in Nightmare mode. |