Hey, wait a minute: on at-issue sensitivity in Language Models (2026.eacl-short)
Copied to clipboard
| Challenge: | Existing methods to evaluate dialogue naturalness are limited. |
| Approach: | They propose a method to assess dialogue naturalness using linguistic notion of at-issueness. |
| Outcome: | The proposed method mitigates bias in linguistic analyses of LMs and tests discourse-sensitive behavior. |
Similar Papers
“No, They Did Not”: Dialogue Response Dynamics in Pre-trained Language Models (2022.coling-1)
Copied to clipboard
| Challenge: | a critical component of competence in language is being able to identify relevant components of an utterance . sensitivity to at-issueness and ellipsis is examined in pre-trained language models . authors find mixed results with respect to capturing full range of dynamics involved in targeting at- issue content . |
| Approach: | They examine sensitivity to at-issueness and ellipsis in pre-trained language models . they find that models show a preference for responses that target main clause content . |
| Outcome: | The proposed models show strong sensitivity to at-issueness and ellipsis dynamics . the results show that the models lack grasp of the dynamics involved in targeting at- issue versus not-at-issue content . |
Behavior-SD: Behaviorally Aware Spoken Dialogue Generation with Large Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Spoken dialogues lack explicit modeling of behavior traits that are often overlooked in language models . et al.: our work opens new possibilities for developing behaviorally-aware dialogue systems . |
| Approach: | They propose a large-scale dataset with over 100K spoken dialogues (2,164 hours) they propose BeDLM, the first dialogue model capable of generating natural conversations . |
| Outcome: | The proposed model outperforms baseline models in generating natural dialogues . the proposed model can generate natural conversations conditioned on behavioral and narrative contexts - a key feature of spoken language models . |
Beyond Static Synthetic Noise: Assessing the Robustness of Large Language Models to Natural Context Variation in the Real World (2026.findings-acl)
Copied to clipboard
| Challenge: | Current robustness evaluation methods rely on static synthetic perturbations to stress-test models. |
| Approach: | They propose a framework for automatically evaluating QA models under naturally occurring textual perturbations by replacing context passages with revised Wikipedia edit histories. |
| Outcome: | The proposed framework replaces context passages with revised Wikipedia edit histories to improve model performance. |
Beyond Prompt Engineering: A Systematic Analysis of Prompt Lexical Sensitivity and Its Impacts on Quality (2026.findings-acl)
Copied to clipboard
Qipeng Xie, Zi Liang, Jiafei Wu, Yufei Chen, Weizheng Wang, Wenao Ma, Zhong Ming, Haiqin Yang, Kaishun Wu
| Challenge: | Existing studies on prompt engineering have focused on optimizing models for performance under stylistic perturbations. |
| Approach: | They conduct the first analysis of n-gram token-level mechanisms . they find that higher average performance is inherently associated with lower variance and greater stability. |
| Outcome: | The proposed model reduces the variance of the generated code by 40% . the proposed model is based on a large-scale dataset of 132,000 prompt variants . |
SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness (2026.acl-long)
Copied to clipboard
Jingyu Lu, Yuhan Wang, Fan Zhuo, Xize Cheng, Changhao Pan, Xueyi Pu, Yifu Chen, Chenyuhao Wen, Tianle Liang, Zhou Zhao
| Challenge: | SDiaReward is an end-to-end spoken dialogue system that integrates paralinguistic nuances and spontaneous nature of human conversation. |
| Approach: | They propose an end-to-end multi-turn reward model trained on SDiaReward-Dataset . it is a collection of episode-level preference pairs targeting modality and colloquiality gaps . |
| Outcome: | The proposed model outperforms general-purpose audio LLMs in episode-level evaluation. |
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are sensitive to subtle, non-semantic variations in prompt phrasing and formatting. |
| Approach: | They propose to evaluate 4 methods for improving prompt robustness within a unified experimental framework. |
| Outcome: | The proposed methods are compared to 8 models from Llama, Qwen and Gemma families and are generalized against multiple types of distribution shifts. |
Automated Fact-Checking in Dialogue: Are Specialized Models Needed? (2023.emnlp-main)
Copied to clipboard
| Challenge: | Prior work has shown that typical fact-checking models struggle with claims made in conversation. |
| Approach: | They propose to fine-tune models for dialogue on conversational data to improve performance on typical fact-checking. |
| Outcome: | The proposed models perform better on stand-alone claims than state-of-the-art models for dialogue while maintaining their performance on standalone claim. |
Lexical Entrainment for Conversational Systems (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Conversational agents are expected to possess human-like features such as lexical entrainment (LE). |
| Approach: | They propose a dataset and a measure for LE for conversational systems to explicitly integrate LE into conversational system. |
| Outcome: | The proposed dataset and a measure for LE for conversational systems address this human-like phenomenon. |
Computational Analysis of Conversation Dynamics through Participant Responsivity (2025.emnlp-main)
Copied to clipboard
| Challenge: | Growing literature explores toxicity and polarization in discourse, with comparatively little work on characterizing what makes dialogue prosocial and constructive. |
| Approach: | They develop and evaluate methods for quantifying responsivity through semantic similarity of speaker turns and large language models to identify the relation between two speaker turns. |
| Outcome: | The proposed method is based on semantic similarity of speaker turns and large language models to identify the relation between two speaker turns. |
Neural Generation of Dialogue Response Timings (2020.acl-main)
Copied to clipboard
| Challenge: | Using neural models, the timings of spoken response offsets in human dialogue can vary based on contextual elements of the dialogue. |
| Approach: | They propose neural models that simulate the distributions of response offsets taking into account the response turn as well as the preceding turn. |
| Outcome: | The proposed models can generate distributions of response offsets based on the response turn and preceding turn based upon human listening tests and offline experiments. |