Challenge: a lot of research aims to mitigate these problems by introducing specific computational solutions.
Approach: They examine how large language models engage in the interactive process of repair in multi-turn dialogues around solvable and unsolvable math questions.
Outcome: The models respond to user-initiated repair differently from one another . the models exhibit their own characteristic form of unreliability in the context of repair .

Similar Papers

Repairs in a Block World: A New Benchmark for Handling User Corrections with Multi-Modal Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: In dialogue, the addressee may misunderstand the speaker and respond erroneously.
Approach: They collect, analyse, and publicly release a dataset of multi-modal TPR sequences in dialogue . they evaluate several state-of-the-art Vision and Language Models across multiple settings .
Outcome: The proposed model underperforms in a human-robot interaction task compared to humans . the proposed model can benefit from specialised losses targeting relevant tokens .
Vulnerability of LLMs’ Stated Belief? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly employed in question-answering tasks.
Approach: They analyze how different persuasive strategies influence stated belief stability . they also examine whether verbalized confidence prompting increases vulnerability .
Outcome: The proposed model exhibits extreme compliance, with 82.5% of belief changes occurring at the first persuasive turn.
Can Language Models Follow Multiple Turns of Entangled Instructions? (2025.findings-emnlp)

Copied to clipboard

Challenge: Despite of significant achievements in improving instruction-following capabilities of large language models, the ability to process multiple potentially entangled or conflicting instructions remains a considerable challenge.
Approach: They construct multi-turn instruction with 1.1K high-quality multi-turned conversations using the human-in-the-loop approach and examine their capabilities.
Outcome: The proposed model shows that it is difficult to integrate multiple turns and balance competing objectives when instructions intersect or conflict.
“Mm, Wat?” Detecting Other-initiated Repair Requests in Dialogue (2025.emnlp-main)

Copied to clipboard

Challenge: Current conversational agents (CAs) do not recognize repair initiation, leading to breakdowns or disengagement.
Approach: They propose a multimodal model to automatically detect repair initiation in Dutch dialogues by integrating linguistic and prosodic features grounded in Conversation Analysis.
Outcome: The proposed model integrates linguistic and prosodic features grounded in Conversation Analysis to detect repair initiation in Dutch dialogues.
One Battle After Another: Probing LLMs’ Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for instruction-following in multi-topic dialogues are limited to a fixed number of turns, susceptible to saturation and failing to account for users’ interactive experience.
Approach: They propose a framework featuring a three-layer tracking mechanism and a query synthesis agent to mimic sequential user behaviors.
Outcome: The proposed framework outperforms existing benchmarks in the evaluation of instruction following in multi-topic dialogues and demonstrates deficiencies in failure recovery and fine-grained instruction following.
LLM as a Broken Telephone: Iterative Generation Distorts Information (2025.acl-long)

Copied to clipboard

Challenge: Large language models are increasingly responsible for online content, but they can be distorted by repeated transmission.
Approach: They investigate whether large language models distort information through iterative generation.
Outcome: The findings raise important questions about the reliability of LLM-generated content in iterative workflows.
Confidence Estimation for LLMs in Multi-turn Interactions (2026.findings-acl)

Copied to clipboard

Challenge: Despite recent progress, most prior work studies confidence in single-turn question answering.
Approach: They propose a logit-based probe that measures confidence in multi-turn dialogues . they propose 'infoECE' and a "hinter-guesser" paradigm for generating controlled evaluations based on data .
Outcome: The proposed framework is grounded in calibration and monotonicity of confidence as more information becomes available.
Human Alignment: How Much Do We Adapt to LLMs? (2025.acl-short)

Copied to clipboard

Challenge: Large Language Models (LLMs) are becoming a common part of our lives, yet few studies have examined how they influence our behavior.
Approach: They propose a cooperative language game in which players aim to converge on a word and play a game in a group.
Outcome: The proposed game shows that humans notice and adapt to differences regardless of whether they are aware they are interacting with an LLM.
Understanding the Dark Side of LLMs’ Intrinsic Self-Correction (2025.acl-long)

Copied to clipboard

Challenge: Recent studies show that LLMs’ intrinsic self-correction fails without oracle labels as feedback.
Approach: They propose to use one simple task and three complex tasks with state-of-the-art LLMs like ChatGPT, Llama, and DeepSeek to interpret LLM's intrinsic self-correction.
Outcome: The proposed methods reveal the dark side of LLMs’ intrinsic self-correction for different tasks, especially for those failure cases.
Persuasion Dynamics in LLMs: Investigating Robustness and Adaptability in Knowledge and Safety with DuET-PD (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) can struggle to balance gullibility to misinformation and resistance to valid corrections in persuasive dialogues.
Approach: They propose a framework evaluating multi-turn stance-change dynamics across dual dimensions: persuasion type and domain.
Outcome: The proposed framework improves LLM-3.1-8B-Instruct accuracy under misleading persuasion in safety contexts from 4.21% to 76.54%.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations