Challenge: Large language models (LLMs) are often prompted to produce natural language explanations, but the features driving the answer are often different from those emphasized in their explanations.
Approach: They propose a large-scale benchmark linking model decisions with diverse explanations and attribution vectors across datasets, methods, and model families to address this gap.
Outcome: The proposed model generates an answer where the word NLP in the prompt has high feature importance.

Similar Papers

Are self-explanations from Large Language Models faithful? (2024.findings-acl)

Copied to clipboard

Challenge: Instruction-tuned Large Language Models excel at many tasks and will explain their reasoning, so-called self-explanations.
Approach: They propose to employ self-consistency checks to measure faithfulness to LLMs to determine if they are model-dependent and if their reasoning is convincing and wrong.
Outcome: The proposed measures show that self-explanations are explanation, model, and task-dependent and should not be trusted in general.
InfoPO: On Mutual Information Maximization for Large Language Model Alignment (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies have shown that direct preference optimization and its variants can be useful for fine-tuning large language models with human preferences data.
Approach: They propose a preference fine-tuning algorithm that effectively and efficiently aligns large language models using preference data.
Outcome: Extensive experiments show that the proposed algorithm outperforms established baselines on reasoning tasks.
On Measuring Faithfulness or Self-consistency of Natural Language Explanations (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) can explain their predictions through post-hoc or Chain-of-Thought explanations, but they can also make unfaithful explanations that hide their sensitivity to biasing inputs.
Approach: They propose to use a model-based consistency test to judge the faithfulness of post-hoc or Chain-of-Thought explanations rather than model-specific faithfulness tests.
Outcome: The proposed tests do not measure faithfulness to model explainability but rather their self-consistency at output level.
From Distributional to Overton Pluralism: Investigating Large Language Model Alignment (2025.naacl-long)

Copied to clipboard

Challenge: a large language model's (LLM) output distribution is changed by an alignment process . a recent study shows that aligned models surface information that cannot be recovered from base models without fine-tuning.
Approach: They analyze two aspects of the alignment process that change output distributions . they find alignment suppresses irrelevant and unhelpful content .
Outcome: The proposed model can be imitated without fine-tuning by using in-context examples and lower-resolution semantic hints about response content.
Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance (2026.acl-long)

Copied to clipboard

Challenge: Prior work has focused on generating convincing rationales that appear to be subjectively faithful, but it remains unclear whether these explanations are epistemic faithful.
Approach: They propose a method that enhances epistemic faithfulness by guiding explanation generation through attention-level interventions, informed by token-level heatmaps.
Outcome: The proposed method significantly improves epistemic faithfulness across multiple models, benchmarks, and prompts.
Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks (2025.acl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) excel in math reasoning problemsolving, text generation, summarization, creative writing, among other tasks.
Approach: They evaluate Direct Preference Optimization and its variants for aligning Large Language Models with human preferences.
Outcome: The proposed alignment methods achieve near-optimal performance even with smaller subsets of training data.
Probability-Consistent Preference Optimization for Enhanced LLM Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in preference optimization have demonstrated significant potential for improving mathematical reasoning capabilities in large language models.
Approach: They propose a framework that establishes two quantitative metrics for preference selection: surface-level answer correctness and intrinsic token-level probability consistency.
Outcome: The proposed framework outperforms existing outcome-only criterion approaches across a diverse range of LLMs and benchmarks.
Rethinking the Evaluation of Alignment Methods: Insights into Diversity, Generalisation, and Safety (2026.eacl-srw)

Copied to clipboard

Challenge: Existing studies focus on individual techniques or specific dimensions, lacking a holistic assessment of the inherent trade-offs.
Approach: They propose a framework that compares LLM alignment methods across five axes . they use a validated LLM-as-judge prompt to compare the results .
Outcome: The proposed framework compares LLM alignment methods across factuality, safety, conciseness, proactivity, diversity and safety axes . it provides insights into trade-offs of common alignment methods, guiding the development of more balanced and reliable LLMs.
Bridging Internal Consistency and External Alignment: A Causal and Dynamic Interpretability Framework for LLM Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing interpretability methods focus on internal and external aspects of the model . existing explanations often focus on surface correlations or static dependencies .
Approach: They propose a causal and dynamic interpretability framework for Large Language Models . they characterize backdoor-adjusted causal effects of generated prefix and prompt .
Outcome: The proposed framework provides a unified causal view of internal consistency and external alignment in LLM generation dynamics.
Uncovering Factor-Level Preference to Improve Human-Model Alignment (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models exhibit tendencies that diverge from human preferences, such as favoring certain writing styles or producing overly verbose outputs.
Approach: They propose a framework to uncover and measure factor-level preference alignment of humans and large language models (LLMs)
Outcome: The proposed framework uncovers and measures factor-level preference alignment of humans and large language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations