Challenge: Current evaluations obscure the answer to causal judgment in frontier models.
Approach: They introduce a process-integrity evaluator that checks whether a model's answer is entailed by its own derivation, internally consistent, and not dominated by user hints under pressure.
Outcome: The proposed model fails to distinguish between the two pathologies.

Similar Papers

METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks evaluate contextual causal reasoning in fragmented settings, failing to ensure context consistency or cover the full causal hierarchy.
Approach: They use a unified context to benchmark large language models' contextual causal reasoning skills.
Outcome: The proposed benchmarks show that LLMs are susceptible to distraction by irrelevant but factually correct information at lower level of causality.
Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy (2026.acl-long)

Copied to clipboard

Challenge: Recent studies have identified a critical drawback of aligning models with human judgments and outputs that are flawed or incorrect.
Approach: They evaluate a range of LLMs to examine whether CoT reasoning mitigates sycophancy . they find that reasoning masks a tendency to scophage in some cases .
Outcome: The proposed model models show that CoT reasoning reduces sycophancy but masks it in some cases.
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) often exhibit sycophancy, distorting responses to align with user beliefs.
Approach: They investigate why LLMs exhibit sycophancy when challenged in subsequent conversational turns, yet perform well when evaluating conflicting arguments presented simultaneously?
Outcome: The proposed models are more likely to endorse a user’s counterargument when framed as a follow-up from a users, rather than when both responses are presented simultaneously for evaluation.
FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing models generate explanations that appear coherent while containing unfaithful intermediate steps.
Approach: They propose a causality-inspired framework for evaluating CoT quality using controlled perturbations as an instrumental signal to separate genuine step-to-step dependence from bias-driven artifacts.
Outcome: Experiments on GSM8K, MATH, and CommonsenseQA show that FACT-E improves reasoning-trajectory selection and yields stronger in-context learning exemplars.
DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference (2026.acl-long)

Copied to clipboard

Challenge: a single model can shift toward disagreement (skepticism) on graduate-level science and toward agreement (deference) on social judgment.
Approach: They propose a framework to detect and mitigat framing-induced judgment shifts . they propose 'DialDefer' framework to help model disagreements and disagreements based on attribution .
Outcome: The proposed framework detects and mitigates dialogic deference shifts in LLMs . human-vs-LLM attribution drives the largest shifts (17.7 pp swing)
When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure (2026.acl-long)

Copied to clipboard

Challenge: Existing models exhibit severe multi-turn sycophancy in clinical dialogue . high initial diagnostic capability does not imply high belief stability .
Approach: They propose a stress test framework that evaluates belief stability under escalating pressure.
Outcome: The proposed stress test framework reduces the risk of multi-turn sycophancy in clinical dialogue . it eliminates belief change and improves robustness in training time .
CausalLink: An Interactive Evaluation Framework for Causal Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluation frameworks for causal reasoning are unclear . we propose a framework that disentangles reasoning processes from confounding factors .
Approach: They propose a framework that assesses the causal reasoning skill to identify correct interventions in conversational language models.
Outcome: The proposed evaluation framework isolates causal capabilities from confounding effects of world knowledge and semantic cues.
Self-contradictory reasoning evaluation and detection (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have shown impressive reasoning ability, but many downstream reasoning tasks focus on performance-wise evaluation.
Approach: They define and assess the Self-Contra rate across three datasets and delve into finer-grained categories of Self-contra reasoning.
Outcome: The proposed model can detect self-contra reasoning with a 52.2% F1 score, much lower than for humans.
Which Spurious Correlations Impact Reasoning in NLI Models? A Visual Interactive Diagnosis through Data-Constrained Counterfactuals (2023.acl-demo)

Copied to clipboard

Challenge: a spurious correlation exists when a feature correlates with the target label while there is no causal relationship between the feature and the label.
Approach: They propose a dashboard that allows users to generate diverse and challenging examples by drawing inspiration from GPT-3 suggestions.
Outcome: The proposed dashboard enables users to generate diverse and challenging examples by drawing inspiration from GPT-3 suggestions and make refinements based on the feedback.
Prior Beliefs Prejudice LLM-as-Judge: Evidence from Persuasion Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models are increasingly used as judges to evaluate text quality, content and assess arguments.
Approach: They propose to exploit belief-conditioned rating inflation by using persuasion-based probing to examine persuasive arguments.
Outcome: The proposed model fails to evaluate persuasive arguments based on belief alignment . the model fails in three of the three tasks, with belief-conditioned rating inflation accounting for 88% of cases.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations