Challenge: Existing models face challenges when dealing with uncertainty.
Approach: They propose a framework that decomposes concept uncertainty by construction . epistemic uncertainty is positively associated with prediction errors, whereas aleatoric uncertainty closely tracks disagreement .
Outcome: The proposed framework decomposes concept uncertainty by construction . epistemic uncertainty is positively associated with prediction errors, whereas aleatoric uncertainty closely tracks disagreement .

Similar Papers

Quantifying Aleatoric Uncertainty of In-Context Learning for Robust Measure of LLM Prediction Confidence (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for standard generation tasks fail to capture the unique dynamics of ICL.
Approach: They propose a concept of self-function vectors that leverage Bayesian views and the mechanistic interpretability of ICL to model latent concept learned during in-context prompting.
Outcome: The proposed framework can be used for trustworthy-related applications, such as hallucination detection.
Do not Abstain! Identify and Solve the Uncertainty (2025.acl-long)

Copied to clipboard

Challenge: Existing solutions rely on evasive responses when confronting uncertain scenarios.
Approach: They propose a benchmark to assess LLMs' ability to recognize and address uncertainty . they generate context-aware inquiries that highlight the confusing aspect of the original query .
Outcome: Experiments with ConfuseBench show that LLMs struggle to identify root cause of uncertainty and solve it.
Beyond "I Don’t Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty (2026.acl-long)

Copied to clipboard

Challenge: Prior studies treat refusal as a generic "I don't know" lack of distinction limits downstream action decisions like requesting clarification or invoking external tools.
Approach: They propose a benchmark to evaluate explicit uncertainty attribution in large language models.
Outcome: The proposed method improves uncertainty attribution while preserving answer accuracy.
Uncertainty Quantification for In-Context Learning of Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Existing studies on in-context learning have focused on quantifying the uncertainty associated with the model's response, but they neglect the complexity of the LLM and the uniqueness of in-constitut learning.
Approach: They propose a method to quantify the uncertainty associated with in-context learning and propose corresponding estimation method to quantify both types of uncertainties.
Outcome: The proposed method offers an unsupervised way to understand the prediction of in-context learning in a plug-and-play fashion.
Demystifying Uncertainty in LLMs: Active Calibration between Concepts and Human Evaluations (2026.acl-long)

Copied to clipboard

Challenge: Existing static strategies for mitigating hallucinations do not explicitly model the information gain from interacting with the external environment.
Approach: They propose a calibration-driven interactive learning strategy that selects clarification queries by optimizing calibration error.
Outcome: The proposed method provides theoretical guarantees and empirical gains for reliability.
Estimating the Black-box LLM Uncertainty with Distribution-Aligned Adversarial Distillation (2026.acl-long)

Copied to clipboard

Challenge: Existing uncertainty quantification methods depend on computationally expensive multiple sampling or internal parameters, which prevents real-time estimation and fails to capture information implicit in the black-box reasoning process.
Approach: They propose a distribution-aligned adjudication architecture to guide a lightweight proxy model to learn the high-quality regions of the output distribution of the black-box LLM.
Outcome: Extensive experiments show that a proxy model even with 1% of the target LLM’s size can achieve reliable uncertainty quantification.
SPUQ: Perturbation-Based Uncertainty Quantification for Large Language Models (2024.eacl-long)

Copied to clipboard

Challenge: Large language models have a tendency to make confidently wrong predictions, highlighting the need for uncertainty quantification (UQ) . previous studies focused on aleatoric uncertainty, but the full spectrum of uncertainties, including epistemic, remains inadequately explored.
Approach: They propose a method to quantify uncertainty in large language models (LLMs) they use a set of perturbations and an aggregation module to generalize the method.
Outcome: The proposed method improves model uncertainty calibration and reduces expected calibration error by 50% on average.
Simple Yet Effective: An Information-Theoretic Approach to Multi-LLM Uncertainty Quantification (2025.emnlp-main)

Copied to clipboard

Challenge: Prior work on calibration and uncertainty quantification focuses on individual models, overlooking the potential of model diversity.
Approach: They propose a method that uses Jensen-Shannon Divergence to identify and aggregate well-calibrated subsets of large language models (LLMs) to improve calibration.
Outcome: The proposed method improves accuracy on binary prediction tasks compared to single-model and naive ensemble baselines.
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding (2026.findings-acl)

Copied to clipboard

Challenge: Autoregressive (AR) decoding in large language models is latency-bounded by strictly sequential token generation.
Approach: They propose a diffusion-based drafter that proposes multi-token candidates and then verifies them in parallel by the target model.
Outcome: The proposed drafter generates multi-token proposals in a single forward pass while remaining compatible with standard AR verifiers.
An Empirical Study of Speculative Decoding for Small Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies focus on 7B-70B parameters models, leaving a knowledge gap for small language models.
Approach: They propose a draft-then-verify paradigm that allows for a single forward pass through a model and transfer of all model parameters to the GPU cache.
Outcome: The proposed method can be used to accelerate small language models with low computational overhead.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations