Challenge: Existing methods to evaluate dialogue naturalness are limited.
Approach: They propose a method to assess dialogue naturalness using linguistic notion of at-issueness.
Outcome: The proposed method mitigates bias in linguistic analyses of LMs and tests discourse-sensitive behavior.

Similar Papers

“No, They Did Not”: Dialogue Response Dynamics in Pre-trained Language Models (2022.coling-1)

Copied to clipboard

Challenge: a critical component of competence in language is being able to identify relevant components of an utterance . sensitivity to at-issueness and ellipsis is examined in pre-trained language models . authors find mixed results with respect to capturing full range of dynamics involved in targeting at- issue content .
Approach: They examine sensitivity to at-issueness and ellipsis in pre-trained language models . they find that models show a preference for responses that target main clause content .
Outcome: The proposed models show strong sensitivity to at-issueness and ellipsis dynamics . the results show that the models lack grasp of the dynamics involved in targeting at- issue versus not-at-issue content .
Behavior-SD: Behaviorally Aware Spoken Dialogue Generation with Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Spoken dialogues lack explicit modeling of behavior traits that are often overlooked in language models . et al.: our work opens new possibilities for developing behaviorally-aware dialogue systems .
Approach: They propose a large-scale dataset with over 100K spoken dialogues (2,164 hours) they propose BeDLM, the first dialogue model capable of generating natural conversations .
Outcome: The proposed model outperforms baseline models in generating natural dialogues . the proposed model can generate natural conversations conditioned on behavioral and narrative contexts - a key feature of spoken language models .
Beyond Static Synthetic Noise: Assessing the Robustness of Large Language Models to Natural Context Variation in the Real World (2026.findings-acl)

Copied to clipboard

Challenge: Current robustness evaluation methods rely on static synthetic perturbations to stress-test models.
Approach: They propose a framework for automatically evaluating QA models under naturally occurring textual perturbations by replacing context passages with revised Wikipedia edit histories.
Outcome: The proposed framework replaces context passages with revised Wikipedia edit histories to improve model performance.
Beyond Prompt Engineering: A Systematic Analysis of Prompt Lexical Sensitivity and Its Impacts on Quality (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on prompt engineering have focused on optimizing models for performance under stylistic perturbations.
Approach: They conduct the first analysis of n-gram token-level mechanisms . they find that higher average performance is inherently associated with lower variance and greater stability.
Outcome: The proposed model reduces the variance of the generated code by 40% . the proposed model is based on a large-scale dataset of 132,000 prompt variants .
SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness (2026.acl-long)

Copied to clipboard

Challenge: SDiaReward is an end-to-end spoken dialogue system that integrates paralinguistic nuances and spontaneous nature of human conversation.
Approach: They propose an end-to-end multi-turn reward model trained on SDiaReward-Dataset . it is a collection of episode-level preference pairs targeting modality and colloquiality gaps .
Outcome: The proposed model outperforms general-purpose audio LLMs in episode-level evaluation.
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are sensitive to subtle, non-semantic variations in prompt phrasing and formatting.
Approach: They propose to evaluate 4 methods for improving prompt robustness within a unified experimental framework.
Outcome: The proposed methods are compared to 8 models from Llama, Qwen and Gemma families and are generalized against multiple types of distribution shifts.
Automated Fact-Checking in Dialogue: Are Specialized Models Needed? (2023.emnlp-main)

Copied to clipboard

Challenge: Prior work has shown that typical fact-checking models struggle with claims made in conversation.
Approach: They propose to fine-tune models for dialogue on conversational data to improve performance on typical fact-checking.
Outcome: The proposed models perform better on stand-alone claims than state-of-the-art models for dialogue while maintaining their performance on standalone claim.
Lexical Entrainment for Conversational Systems (2023.findings-emnlp)

Copied to clipboard

Challenge: Conversational agents are expected to possess human-like features such as lexical entrainment (LE).
Approach: They propose a dataset and a measure for LE for conversational systems to explicitly integrate LE into conversational system.
Outcome: The proposed dataset and a measure for LE for conversational systems address this human-like phenomenon.
Computational Analysis of Conversation Dynamics through Participant Responsivity (2025.emnlp-main)

Copied to clipboard

Challenge: Growing literature explores toxicity and polarization in discourse, with comparatively little work on characterizing what makes dialogue prosocial and constructive.
Approach: They develop and evaluate methods for quantifying responsivity through semantic similarity of speaker turns and large language models to identify the relation between two speaker turns.
Outcome: The proposed method is based on semantic similarity of speaker turns and large language models to identify the relation between two speaker turns.
Neural Generation of Dialogue Response Timings (2020.acl-main)

Copied to clipboard

Challenge: Using neural models, the timings of spoken response offsets in human dialogue can vary based on contextual elements of the dialogue.
Approach: They propose neural models that simulate the distributions of response offsets taking into account the response turn as well as the preceding turn.
Outcome: The proposed models can generate distributions of response offsets based on the response turn and preceding turn based upon human listening tests and offline experiments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations