Accounting for Sycophancy in Language Model Uncertainty Estimation (2025.findings-naacl)
Copied to clipboard
| Challenge: | Effective human-machine collaboration requires machine learning models to externalize uncertainty. |
| Approach: | They propose a generalization of the definition of sycophancy bias and a new algorithm to account for scophancies in uncertainty estimation. |
| Outcome: | The proposed algorithm can account for sycophancy in uncertainty estimation process. |
Similar Papers
Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models are increasingly used as conversational agents that adopt personas and role-play characters at user request. |
| Approach: | They propose to examine how persona agreeableness influences sycophancy across 13 small, open-weight language models ranging from 0.6B to 20B parameters. |
| Outcome: | The proposed model consists of 275 personas and exposes them to 4,950 sycophancy-eliciting prompts spanning 33 topic categories. |
Measuring Sycophancy of Language Models in Multi-turn Dialogues (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Prior research on sycophancy has focused on single-turn factual correctness, overlooking the dynamics of real-world interactions. |
| Approach: | They propose a new evaluation suite that assesses sycophantic behavior in multi-turn, free-form conversational settings. |
| Outcome: | The proposed evaluation suite measures how quickly a model conforms to the user and how frequently it shifts its stance under sustained user pressure. |
Echoes of Agreement: Argument Driven Sycophancy in Large Language models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluations of political biases in Large Language Models outline the high sensitivity to prompt formulation. |
| Approach: | They investigate how argumentative prompts induce sycophantic behaviour in Large Language Models in a political context. |
| Outcome: | The proposed model sycophancy is observed in single and multi-turn interactions and its intensity correlates with argument strength. |
Perceptions of Linguistic Uncertainty by Language Models and Humans (2024.emnlp-main)
Copied to clipboard
| Challenge: | Prior work has shown that humans are well-attuned to the use of uncertainty expressions, exhibiting population-level agreement in mapping these expressions to numerical responses. |
| Approach: | They propose to map linguistic expressions of uncertainty to numerical responses by using a theory of mind approach to understand the uncertainty of another agent. |
| Outcome: | The proposed model can map expressions to probabilistic responses in a human-like manner, but different behavior depending on whether a statement is actually true or false. |
Self-Augmented Preference Alignment for Sycophancy Reduction in LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Sycophantic behavior in models can erode user trust by creating a perception of dishonesty or bias. |
| Approach: | They propose to assess the user’s expected answer rather than ignore it and introduce self-augmented preference alignment to reduce sycophancy. |
| Outcome: | The proposed methods significantly reduce sycophancy across tasks and improve models' assessment ability. |
Learning Multilingual Agentic Policy to Control Sycophancy (2026.eacl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are effective at adapting to users’ styles, preferences, and contextual signals, but can manifest as sycophancy, i.e., alignment with user-implied beliefs or assumptions even when these contradict factual correctness, uncertainty, or proper logical reasoning. |
| Approach: | They propose to use large language models to model sycophancy as a decision-making problem by learning agentic policies that are trained to optimise a multi-objective reward that balances task success, scophancies resistance and behavioural consistency. |
| Outcome: | The proposed model equips a model with an explicit action space that includes answering directly, countering misleading signals, or asking for clarification. |
Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy (2026.acl-long)
Copied to clipboard
| Challenge: | Recent studies have identified a critical drawback of aligning models with human judgments and outputs that are flawed or incorrect. |
| Approach: | They evaluate a range of LLMs to examine whether CoT reasoning mitigates sycophancy . they find that reasoning masks a tendency to scophage in some cases . |
| Outcome: | The proposed model models show that CoT reasoning reduces sycophancy but masks it in some cases. |
Sounding vs. Being an Expert: Disentangling Authority, Register and Cultural Impact in Sycophantic LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models exhibit sycophancy, a tendency to align with user assertions even when they conflict with factual correctness. |
| Approach: | They propose an adversarial evaluation framework that isolates two drivers of credibility: explicit authority (credentials) and implicit authority (linguistic register). |
| Outcome: | The proposed framework disentangles two drivers of credibility: explicit authority (credentials) and implicit authority (linguistic register). |
Advancing Oversight Reasoning across Languages for Audit Sycophantic Behaviour via X-Agent (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models have demonstrated capabilities that are satisfactory to a wide range of users by adapting to their culture and wisdom. |
| Approach: | They propose an Oversight Reasoning framework that audits human–LLM dialogues, reasons about them, captures sycophancy and corrects the final outputs. |
| Outcome: | The proposed framework detects sycophancy, reduces unwarranted agreement and improves cross-turn consistency across different scenarios and languages. |
A Survey of Uncertainty Estimation Methods on Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated remarkable capabilities but could produce biased, hallucinated, or non-factual responses. |
| Approach: | They propose to conduct extensive experimental evaluations of LLM uncertainty estimation methods . large language models have demonstrated remarkable capabilities across tasks . |
| Outcome: | The proposed method could produce biased, hallucinated, or non-factual responses . a lack of comprehensive surveys on LLM uncertainty estimation is a problem . |