Challenge: a new study shows that voice-based search systems are challenging to support in the context of the user intent of voice searches . support for voice-driven search, exploration, and refinement is a fundamental aspect of voice assistants .
Approach: They propose to use crowdsourcing to collect voice-based search refinements . they use 10,000 search refinement utterances to annotate a search intent .
Outcome: The proposed dataset shows that voice-based search refinements can support most common tasks . the study shows that the proposed dataset can support research in conversational query understanding .

Similar Papers

VoiceBench: Benchmarking LLM-Based Voice Assistants (2026.tacl-1)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have enabled real-time speech interactions through LLMs.
Approach: They propose a benchmark specifically designed to assess LLM-based voice assistants.
Outcome: The proposed benchmark measures the performance of LLM-based voice assistants across eight tasks.
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for evaluating the perceptual quality of synthetic speech are limited due to the complexity of perceptual quality factors and the diversity of speech generation tasks.
Approach: They propose a new paradigm for enabling large language models to conduct structured speech quality evaluation using a large-scale dataset.
Outcome: The proposed model performs well across tasks and languages.
Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in multi-turn voice interaction models have improved user-model communication, but whether open-source models share this ability remains unexplored.
Approach: They propose to use ContextDialog to evaluate open-source interaction models' ability to recall past utterances to identify key limitations.
Outcome: The proposed model retains and recalls past utterances better than closed-source models, but still struggles with questions about past . findings highlight key limitations in open-source model and suggest ways to improve memory retention and retrieval robustness.
PRESTO: A Multilingual Dataset for Parsing Realistic Task-Oriented Dialogs (2023.emnlp-main)

Copied to clipboard

Challenge: PRESTO dataset contains 550K contextual multilingual conversations between humans and virtual assistants.
Approach: They propose to use a dataset of 550K contextual multilingual conversations between humans and virtual assistants to study some of the more challenging aspects of parsing realistic conversations.
Outcome: The dataset contains 550K contextual conversations between humans and virtual assistants.
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants (2025.findings-acl)

Copied to clipboard

Challenge: Existing personalization benchmarks focus on chit-chat, non-conversational tasks, or narrow domains, failing to capture complexities of personalized task-oriented assistance.
Approach: They propose a benchmark to evaluate personalization in task-oriented AI assistants . the benchmark features user profiles equipped with rich preferences and interaction histories .
Outcome: The proposed benchmark features user profiles equipped with rich preferences and interaction histories . it also features a judge agent and user agent that employs the LLM-as-a-Judge paradigm .
It’s Not under the Lamppost: Expanding the Reach of Conversational AI (2024.lrec-main)

Copied to clipboard

Challenge: Focused probes into the capabilities of language-based assistants easily reveal significant areas of brittleness that demonstrate large gaps in their coverage.
Approach: They propose a process for collecting specific kinds of data to uncover these gaps and an annotation scheme for system responses.
Outcome: The proposed system includes both Conventional and GenAI systems, including ChatGPT and Bard/Gemini.
LingVarBench: Benchmarking LLMs on Entity Recognitions and Linguistic Verbalization Patterns in Phone-Call Transcripts (2026.eacl-industry)

Copied to clipboard

Challenge: Existing methods degrade under disfluencies, interruptions, and speaker overlap, yet large real-call corpora are rarely shareable.
Approach: They propose a benchmark and semantic synthetic data generation pipeline that generates linguistically varied training data via (1) LLM-sampled entity values, (2) curated linguistic verbalization patterns covering diverse disfluencies and entity-specific readout styles, and (3) a value–transcript consistency filter.
Outcome: The proposed pipeline outperforms zero-shot baselines and matches or closely approaches human-tuned prompts on real customer transcripts.
CL-QR: Cross-Lingual Enhanced Query Reformulation for Multi-lingual Conversational AI Agents (2023.emnlp-industry)

Copied to clipboard

Challenge: Existing QR systems that reformulate defective user queries are limited in English due to the scarcity of non-English QR labels.
Approach: They propose a query reformulation method which reformulates defective user queries to improve non-English QR performance.
Outcome: The proposed framework improves non-English QR performance by leveraging abundant reformulation resources in English.
WESR: A Benchmark and Strong Baseline for Word-level Event-Speech Recognition (2026.findings-acl)

Copied to clipboard

Challenge: aaron carroll: the precise localization of non-verbal vocal events remains a critical yet under-explored challenge. carroll says current methods suffer from insufficient task definitions with limited category coverage. carrol: knowing exactly where an event occurred is not enough; knowing exactly what it happened is.
Approach: They propose a taxonomy of 21 vocal events with a new categorization into discrete versus continuous types.
Outcome: The proposed model disentangles ASR errors from event detection while maintaining ASR quality.
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases (2026.eacl-long)

Copied to clipboard

Challenge: Existing models that measure audio comprehension beyond automatic speech recognition lack performance and latency.
Approach: They propose a benchmark suite that measures audio comprehension beyond automatic speech recognition . the benchmark suite includes a small human-recorded evaluation split per category .
Outcome: The proposed suite measures audio comprehension beyond speech recognition . it includes a small human-recorded evaluation split per category .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations