Computational Analysis of Conversation Dynamics through Participant Responsivity (2025.emnlp-main)
Copied to clipboard
| Challenge: | Growing literature explores toxicity and polarization in discourse, with comparatively little work on characterizing what makes dialogue prosocial and constructive. |
| Approach: | They develop and evaluate methods for quantifying responsivity through semantic similarity of speaker turns and large language models to identify the relation between two speaker turns. |
| Outcome: | The proposed method is based on semantic similarity of speaker turns and large language models to identify the relation between two speaker turns. |
Similar Papers
Can Language Model Moderators Improve the Health of Online Discourse? (2024.naacl-long)
Copied to clipboard
Hyundong Cho, Shuai Liu, Taiwei Shi, Darpan Jain, Basem Rizk, Yuyang Huang, Zixun Lu, Nuan Wen, Jonathan Gratch, Emilio Ferrara, Jonathan May
| Challenge: | Existing efforts to automate conversational moderation have focused on banning harmful comments or deleting them, but such efforts can inadvertently push users towards echo chambers that exacerbate polarization. |
| Approach: | They propose a framework to assess models’ moderation capabilities independently of human intervention and propose 'conversational moderation' they propose to use language models as conversational moderators to provide specific feedback on toxic behavior but struggle to influence users to increase their levels of respect and cooperation. |
| Outcome: | The proposed framework assesses models’ moderation capabilities independently of human intervention and shows that appropriately prompted models provide specific and fair feedback on toxic behavior but struggle to influence users to increase their levels of respect and cooperation. |
A Similarity Measure for Comparing Conversational Dynamics (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Qualities of a conversation are dependent on how interactions combine to form a “shape” of the conversation. |
| Approach: | They propose a similarity measure to capture differences in conversation dynamics and assess its sensitivity to the topic of the conversation. |
| Outcome: | The proposed measure captures differences in conversation dynamics and assesses its sensitivity to the topic of the conversation. |
ProsocialDialog: A Prosocial Backbone for Conversational Agents (2022.emnlp-main)
Copied to clipboard
Hyunwoo Kim, Youngjae Yu, Liwei Jiang, Ximing Lu, Daniel Khashabi, Gunhee Kim, Yejin Choi, Maarten Sap
| Challenge: | Existing dialogue systems fail to respond properly to potentially unsafe user utterances . existing systems either ignore or passively agree with unsafe content . |
| Approach: | They introduce a dataset to teach conversational agents to respond to problematic content following social norms. |
| Outcome: | The proposed dataset shows that ProsocialDialog generates more socially acceptable dialogues than existing models. |
Characterizing Similarities and Divergences in Conversational Tones in Humans and LLMs by Sampling with People (2024.acl-long)
Copied to clipboard
| Challenge: | Existing taxonomies or text corpora suffer from experimenter bias and are not representative of real-world distributions. |
| Approach: | They propose an iterative method for simultaneously eliciting conversational tones and sentences . they run 50 iterations with human participants and GPT-4 and obtain a dataset of sentences and frequent conversational tone. |
| Outcome: | The proposed method can be used to characterize the differences between humans and LLMs. |
RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language Models (2021.acl-long)
Copied to clipboard
| Challenge: | Recent work has focused on measuring and mitigating bias in pretrained language models. |
| Approach: | They propose a dataset that measures and mitigates bias across gender,race, religion, and queerness . they compare REDDITBIAS to a widely used conversational DialoGPT model . |
| Outcome: | The proposed framework measures and mitigates bias across gender,race, religion, and queerness dimensions. |
Evaluation and Facilitation of Online Discussions in the LLM Era: A Survey (2025.emnlp-main)
Copied to clipboard
Katerina Korre, Dimitris Tsirmpas, Nikos Gkoumas, Emma Cabalé, Danai Myrtzani, Theodoros Evgeniou, Ion Androutsopoulos, John Pavlopoulos
| Challenge: | Recent advances in LLMs enable artificial facilitation agents to not only moderate content, but also actively improve the quality of interactions. |
| Approach: | They propose a taxonomy on discussion quality evaluation and a new taxonomies for intervention and facilitation strategies. |
| Outcome: | The proposed methods synthesize ideas from Natural Language Processing (NLP) and Social Sciences to provide a taxonomy on discussion quality evaluation, and a roadmap of good practices and future research directions. |
CIG: Measuring Conversational Information Gain in Deliberative Dialogues with Semantic Memory Dynamics (2026.acl-long)
Copied to clipboard
| Challenge: | Using a semantic memory, we score each utterance along three interpretable dimensions: Novelty, Relevance, and Implication Scope. |
| Approach: | They propose a framework for Conversational Information Gain that evaluates each utterance in terms of how it advances collective understanding of the target topic. |
| Outcome: | The proposed framework evaluates each utterance in terms of how it advances collective understanding of the target topic. |
Behavior-SD: Behaviorally Aware Spoken Dialogue Generation with Large Language Models (2025.naacl-long)
Copied to clipboard
| Challenge: | Spoken dialogues lack explicit modeling of behavior traits that are often overlooked in language models . et al.: our work opens new possibilities for developing behaviorally-aware dialogue systems . |
| Approach: | They propose a large-scale dataset with over 100K spoken dialogues (2,164 hours) they propose BeDLM, the first dialogue model capable of generating natural conversations . |
| Outcome: | The proposed model outperforms baseline models in generating natural dialogues . the proposed model can generate natural conversations conditioned on behavioral and narrative contexts - a key feature of spoken language models . |
Do LLM Agents Mirror Socio-Cognitive Effects in Power-Asymmetric Conversations? (2026.acl-long)
Copied to clipboard
| Challenge: | Power differences shape human communication through well-documented socio-cognitive effects . asymmetric relationships or power differentials give rise to well-known socio-computational effects - lianelli, 1976 . |
| Approach: | They simulate multi-turn, power-asymmetric dialogues with personas from diverse professions . they find that LLMs show key socio-cognitive effects of power, albeit with nuances and variability . |
| Outcome: | The results show that large language models exhibit socio-cognitive effects of power . the results are consistent with previous studies on LLMs . |
Probing the Robustness of Trained Metrics for Conversational Dialogue Systems (2022.acl-short)
Copied to clipboard
| Challenge: | Existing methods for evaluating conversational dialogue systems have been shown to be inefficient and instabile. |
| Approach: | They propose an adversarial method to stress-test trained metrics for evaluation of conversational dialogue systems using Reinforcement Learning. |
| Outcome: | The proposed method outperforms existing methods and can be applied to stress-test trained metrics for conversational dialogue systems. |