Challenge: Existing evaluations of large language models' ability to communicate uncertainty and knowledge limitations focus on the behaviors of their human interlocutors.
Approach: They propose an interaction-centered evaluation approach that quantifies whether and how humans rely on LLMs' responses.
Outcome: The proposed approach quantifies whether and how humans rely on LLMs' responses.

Similar Papers

Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks (2025.coling-main)

Copied to clipboard

Challenge: Existing work uses large language models (LLMs) to evaluate natural language process tasks, but there are shortcomings in current LLMs.
Approach: They examine the alignment between LLM evaluators and human annotators by comparing conventional and alignment tasks with different evaluation criteria.
Outcome: The proposed models excel in general criteria, such as fluency, but face challenges with complex criteria, including numerical reasoning.
Relying on the Unreliable: The Impact of Language Models’ Reluctance to Express Uncertainty (2024.acl-long)

Copied to clipboard

Challenge: a pivotal aspect of fostering reliable human-AI interactions lies in the apt communication of model confidences.
Approach: They examine how LMs incorporate confidence in responses via natural language . they also examine how downstream users behave in response to LM-articulated uncertainties .
Outcome: The proposed model overconfidences are high in LMs, and humans are biased against uncertainty-rich texts.
Demystifying Uncertainty in LLMs: Active Calibration between Concepts and Human Evaluations (2026.acl-long)

Copied to clipboard

Challenge: Existing static strategies for mitigating hallucinations do not explicitly model the information gain from interacting with the external environment.
Approach: They propose a calibration-driven interactive learning strategy that selects clarification queries by optimizing calibration error.
Outcome: The proposed method provides theoretical guarantees and empirical gains for reliability.
A Survey of Confidence Estimation and Calibration in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated impressive capabilities across a wide range of tasks in various domains, but they can be unreliable due to factual errors in their generations.
Approach: They summarize recent advances in LLM confidence estimation and calibration and outline their main lessons learned.
Outcome: The proposed methods can be used to assess the reliability of models and to calibrate them across tasks.
Calibrating Long-form Generations From Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Conventional calibration methods treat answer correctness as binary and do not work for long-form generation where an answer can be partially correct.
Approach: They propose a framework where correctness of LLMs' responses and associated confidence levels are treated as distributions across a range of scores.
Outcome: The proposed framework treats the correctness of the LLMs’ responses and their associated confidence levels as distributions across a range of scores.
Learning to Judge: LLMs Designing and Applying Evaluation Rubrics (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models are increasingly used as evaluators for natural language generation . human rubrics are often static and misaligned with how models internally represent language quality.
Approach: They propose to use large language models to generate interpretable and task-aware evaluation dimensions and apply them within models.
Outcome: The proposed model improves the semantic coherence and scoring reliability of LLM-defined criteria and their alignment with human criteria.
SaGE: Evaluating Moral Consistency in Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on Large Language Models (LLMs) have focused on accuracy but lack universally agreed-upon answers for moral scenarios.
Approach: They propose a measure called Semantic Graph Entropy to measure a model's moral consistency grounded in "Rules of Thumb" they construct a moral Consistency Corpus (MCC) with 50K moral questions and the RoTs they followed to investigate LLM consistency on two popular datasets.
Outcome: The proposed measure measures moral consistency on two popular datasets .
Social Intelligence in the Age of LLMs (2025.naacl-tutorial)

Copied to clipboard

Challenge: Large Language Models (LLMs) are a powerful tool for integrating human-like communication and context-aware interactions into artificial systems.
Approach: They propose to introduce and overview different aspects of artificial social intelligence and their relationship with LLMs by introducing scientific methods for evaluating social intelligence in LLM.
Outcome: This tutorial will introduce scientific methods for evaluating social intelligence in LLMs, highlighting the key challenges, and identifying promising research directions.
Can a Large Language Model Keep My Secrets? A Study on LLM-Controlled Agents (2025.acl-srw)

Copied to clipboard

Challenge: Using large language models, agents can assist with natural language tasks when given access to confidential data.
Approach: They created a synthetic dataset consisting of confidentiality-aware planning and deduction tasks in organizational access control.
Outcome: The proposed model can perform tasks similar to humans when given access to confidential data.
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used as automated evaluators . et al., 2024: strong labels can foster trust but also undermine it .
Approach: They show that LLMs' source labels bias trust judgments by humans . they use eye-tracking data to analyze LLM internal states during judgment .
Outcome: The proposed model is biased by disclosed source labels, the authors show . eye-tracking data show humans rely heavily on source labels for judgments .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations