Sounding vs. Being an Expert: Disentangling Authority, Register and Cultural Impact in Sycophantic LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models exhibit sycophancy, a tendency to align with user assertions even when they conflict with factual correctness. |
| Approach: | They propose an adversarial evaluation framework that isolates two drivers of credibility: explicit authority (credentials) and implicit authority (linguistic register). |
| Outcome: | The proposed framework disentangles two drivers of credibility: explicit authority (credentials) and implicit authority (linguistic register). |
Similar Papers
Echoes of Agreement: Argument Driven Sycophancy in Large Language models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluations of political biases in Large Language Models outline the high sensitivity to prompt formulation. |
| Approach: | They investigate how argumentative prompts induce sycophantic behaviour in Large Language Models in a political context. |
| Outcome: | The proposed model sycophancy is observed in single and multi-turn interactions and its intensity correlates with argument strength. |
Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards? (2026.acl-long)
Copied to clipboard
| Challenge: | Recent advances have seen large language models (LLMs) achieve remarkable performance across high-stakes specialized domains. |
| Approach: | They propose a diagnostic framework that evaluates legal reasoning against medical baselines along four axes (knowledge recall, grounding, confidence, and robustness) they uncover a sharp domain asymmetry when applied to a benchmark that encodes temporal validity and normative relationships. |
| Outcome: | The proposed framework evaluates legal reasoning against medical baselines along four axes (knowledge recall, grounding, confidence, and robustness) it shows that legal LLMs struggle to assess when retrieved citations are useful or misleading, exhibiting overconfidence in perturbed contexts and sensitivity to superficial formatting cues. |
Measuring Sycophancy of Language Models in Multi-turn Dialogues (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Prior research on sycophancy has focused on single-turn factual correctness, overlooking the dynamics of real-world interactions. |
| Approach: | They propose a new evaluation suite that assesses sycophantic behavior in multi-turn, free-form conversational settings. |
| Outcome: | The proposed evaluation suite measures how quickly a model conforms to the user and how frequently it shifts its stance under sustained user pressure. |
Learning Multilingual Agentic Policy to Control Sycophancy (2026.eacl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are effective at adapting to users’ styles, preferences, and contextual signals, but can manifest as sycophancy, i.e., alignment with user-implied beliefs or assumptions even when these contradict factual correctness, uncertainty, or proper logical reasoning. |
| Approach: | They propose to use large language models to model sycophancy as a decision-making problem by learning agentic policies that are trained to optimise a multi-objective reward that balances task success, scophancies resistance and behavioural consistency. |
| Outcome: | The proposed model equips a model with an explicit action space that includes answering directly, countering misleading signals, or asking for clarification. |
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) often exhibit sycophancy, distorting responses to align with user beliefs. |
| Approach: | They investigate why LLMs exhibit sycophancy when challenged in subsequent conversational turns, yet perform well when evaluating conflicting arguments presented simultaneously? |
| Outcome: | The proposed models are more likely to endorse a user’s counterargument when framed as a follow-up from a users, rather than when both responses are presented simultaneously for evaluation. |
Advancing Oversight Reasoning across Languages for Audit Sycophantic Behaviour via X-Agent (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models have demonstrated capabilities that are satisfactory to a wide range of users by adapting to their culture and wisdom. |
| Approach: | They propose an Oversight Reasoning framework that audits human–LLM dialogues, reasons about them, captures sycophancy and corrects the final outputs. |
| Outcome: | The proposed framework detects sycophancy, reduces unwarranted agreement and improves cross-turn consistency across different scenarios and languages. |
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly studied as repositories of linguistic knowledge. |
| Approach: | They compare LLMs’ performance as pragmatic listeners and as pragmatic speakers . they find a robust asymmetry between pragmatic evaluation and pragmatic generation . |
| Outcome: | The proposed models perform better as listeners than speakers, and produce more appropriate language than speakers. |
Bias in the Mirror : Are LLMs opinions robust to their own adversarial attacks (2025.acl-long)
Copied to clipboard
| Challenge: | Existing work on large language models lacks robustness, highlighting the limitations of such models. |
| Approach: | They propose a novel approach where two LLMs engage in self-debate to persuade a neutral version of the model. |
| Outcome: | The proposed approach examines whether large language models are robust during interactions and whether they are susceptible to reinforcing misinformation or shifting to harmful viewpoints. |
Self-Augmented Preference Alignment for Sycophancy Reduction in LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Sycophantic behavior in models can erode user trust by creating a perception of dishonesty or bias. |
| Approach: | They propose to assess the user’s expected answer rather than ignore it and introduce self-augmented preference alignment to reduce sycophancy. |
| Outcome: | The proposed methods significantly reduce sycophancy across tasks and improve models' assessment ability. |
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Current sycophancy research has largely overlooked its specific manifestations in the video-language domain. |
| Approach: | They propose a video-LLM sycophancy benchmarking and evaluation to evaluate scophancies in video-LLMs. |
| Outcome: | The proposed benchmark evaluates sycophantic behavior in state-of-the-art Video-LLMs across diverse question formats, prompt biases, and visual reasoning tasks. |