Challenge: Effective human-machine collaboration requires machine learning models to externalize uncertainty.
Approach: They propose a generalization of the definition of sycophancy bias and a new algorithm to account for scophancies in uncertainty estimation.
Outcome: The proposed algorithm can account for sycophancy in uncertainty estimation process.

Similar Papers

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large language models are increasingly used as conversational agents that adopt personas and role-play characters at user request.
Approach: They propose to examine how persona agreeableness influences sycophancy across 13 small, open-weight language models ranging from 0.6B to 20B parameters.
Outcome: The proposed model consists of 275 personas and exposes them to 4,950 sycophancy-eliciting prompts spanning 33 topic categories.
Measuring Sycophancy of Language Models in Multi-turn Dialogues (2025.findings-emnlp)

Copied to clipboard

Challenge: Prior research on sycophancy has focused on single-turn factual correctness, overlooking the dynamics of real-world interactions.
Approach: They propose a new evaluation suite that assesses sycophantic behavior in multi-turn, free-form conversational settings.
Outcome: The proposed evaluation suite measures how quickly a model conforms to the user and how frequently it shifts its stance under sustained user pressure.
Echoes of Agreement: Argument Driven Sycophancy in Large Language models (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluations of political biases in Large Language Models outline the high sensitivity to prompt formulation.
Approach: They investigate how argumentative prompts induce sycophantic behaviour in Large Language Models in a political context.
Outcome: The proposed model sycophancy is observed in single and multi-turn interactions and its intensity correlates with argument strength.
Perceptions of Linguistic Uncertainty by Language Models and Humans (2024.emnlp-main)

Copied to clipboard

Challenge: Prior work has shown that humans are well-attuned to the use of uncertainty expressions, exhibiting population-level agreement in mapping these expressions to numerical responses.
Approach: They propose to map linguistic expressions of uncertainty to numerical responses by using a theory of mind approach to understand the uncertainty of another agent.
Outcome: The proposed model can map expressions to probabilistic responses in a human-like manner, but different behavior depending on whether a statement is actually true or false.
Self-Augmented Preference Alignment for Sycophancy Reduction in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Sycophantic behavior in models can erode user trust by creating a perception of dishonesty or bias.
Approach: They propose to assess the user’s expected answer rather than ignore it and introduce self-augmented preference alignment to reduce sycophancy.
Outcome: The proposed methods significantly reduce sycophancy across tasks and improve models' assessment ability.
Learning Multilingual Agentic Policy to Control Sycophancy (2026.eacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are effective at adapting to users’ styles, preferences, and contextual signals, but can manifest as sycophancy, i.e., alignment with user-implied beliefs or assumptions even when these contradict factual correctness, uncertainty, or proper logical reasoning.
Approach: They propose to use large language models to model sycophancy as a decision-making problem by learning agentic policies that are trained to optimise a multi-objective reward that balances task success, scophancies resistance and behavioural consistency.
Outcome: The proposed model equips a model with an explicit action space that includes answering directly, countering misleading signals, or asking for clarification.
Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy (2026.acl-long)

Copied to clipboard

Challenge: Recent studies have identified a critical drawback of aligning models with human judgments and outputs that are flawed or incorrect.
Approach: They evaluate a range of LLMs to examine whether CoT reasoning mitigates sycophancy . they find that reasoning masks a tendency to scophage in some cases .
Outcome: The proposed model models show that CoT reasoning reduces sycophancy but masks it in some cases.
Sounding vs. Being an Expert: Disentangling Authority, Register and Cultural Impact in Sycophantic LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models exhibit sycophancy, a tendency to align with user assertions even when they conflict with factual correctness.
Approach: They propose an adversarial evaluation framework that isolates two drivers of credibility: explicit authority (credentials) and implicit authority (linguistic register).
Outcome: The proposed framework disentangles two drivers of credibility: explicit authority (credentials) and implicit authority (linguistic register).
Advancing Oversight Reasoning across Languages for Audit Sycophantic Behaviour via X-Agent (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models have demonstrated capabilities that are satisfactory to a wide range of users by adapting to their culture and wisdom.
Approach: They propose an Oversight Reasoning framework that audits human–LLM dialogues, reasons about them, captures sycophancy and corrects the final outputs.
Outcome: The proposed framework detects sycophancy, reduces unwarranted agreement and improves cross-turn consistency across different scenarios and languages.
A Survey of Uncertainty Estimation Methods on Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable capabilities but could produce biased, hallucinated, or non-factual responses.
Approach: They propose to conduct extensive experimental evaluations of LLM uncertainty estimation methods . large language models have demonstrated remarkable capabilities across tasks .
Outcome: The proposed method could produce biased, hallucinated, or non-factual responses . a lack of comprehensive surveys on LLM uncertainty estimation is a problem .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations