Challenge: Large Language Models (LLMs) generate misleading or outright incorrect information.
Approach: They propose a method that debiases uncertainty scores on output length and uses residuals as corrected, length-invariant estimates.
Outcome: The proposed method improves over nominally length-normalized methods on machine translation, summarization, and question-answering tasks.

Similar Papers

A Survey of Uncertainty Estimation Methods on Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable capabilities but could produce biased, hallucinated, or non-factual responses.
Approach: They propose to conduct extensive experimental evaluations of LLM uncertainty estimation methods . large language models have demonstrated remarkable capabilities across tasks .
Outcome: The proposed method could produce biased, hallucinated, or non-factual responses . a lack of comprehensive surveys on LLM uncertainty estimation is a problem .
LUQ: Long-text Uncertainty Quantification for LLMs (2024.emnlp-main)

Copied to clipboard

Challenge: Existing research on Uncertainty Quantification (UQ) predominantly targets short text generation, however, real-world applications often necessitate much longer responses.
Approach: They propose a method that ensembles responses from multiple models and selects the response with the lowest uncertainty.
Outcome: The proposed method outperforms baseline methods in correlating with the model’s factuality scores (negative coefficient of -0.85 observed for Gemini Pro).
Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) show promising results in language generation but often “hallucinate”, making their outputs less reliable.
Approach: They propose to shift attention to more relevant components at token- and sentence-levels for better UQ.
Outcome: The proposed approach improves the performance of a range of popular “off-the-shelf” LLMs with model sizes extending up to 33B parameters.
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results (2025.acl-short)

Copied to clipboard

Challenge: Language Models (LMs) produce factually incorrect outputs, or "hallucinations" Xiao and Wang et al., 2023) rely on AUROC to assess how well UQ methods distinguish correct from incorrect output.
Approach: They propose to use length biases in correctness functions to skew UQ evaluations . they propose to employ LM-as-a-judge methods as the least length-biased .
Outcome: The proposed method is least length-biased, offering a promising path for a fairer evaluation.
A Survey of Confidence Estimation and Calibration in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated impressive capabilities across a wide range of tasks in various domains, but they can be unreliable due to factual errors in their generations.
Approach: They summarize recent advances in LLM confidence estimation and calibration and outline their main lessons learned.
Outcome: The proposed methods can be used to assess the reliability of models and to calibrate them across tasks.
Towards Harmonized Uncertainty Estimation for Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated exceptional capabilities in handling a wide range of downstream tasks.
Approach: They propose a method that employs a lightweight model trained on data aligned with the target LLM’s performance to adjust uncertainty scores.
Outcome: The proposed method achieves improvements of up to 60% over existing methods.
Same evaluation, more tokens: On the effect of input length for machine translation evaluation using Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models can be used to evaluate long documents, but they are limited by context window limitations.
Approach: They propose to use granularity-aligned prompting and Focus Sentence Prompting to improve evaluation.
Outcome: a new study shows that long texts lead to fewer error spans and reduced system ranking accuracy.
IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model Generation (2026.acl-long)

Copied to clipboard

Challenge: Recent approaches to quantify uncertainty in LLMs produce short or constrained answer sets, but many real-world applications require long-form and free-form text generation.
Approach: They propose a framework that leverages inter-sample consistency and intra-sampled faithfulness to quantify the uncertainty in long-form LLM outputs.
Outcome: The proposed framework provides reliable measures of claim-level uncertainty and the model’s faithfulness over two widely used long-form generation datasets.
Token-Level Density-Based Uncertainty Quantification Methods for Eliciting Truthfulness of Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Uncertainty quantification (UQ) is a prominent approach for eliciting truthful answers from large language models (LLMs).
Approach: They propose to use a well-established method for text generation to extract token embeddings from multiple layers of LLMs and compute MD scores for each token.
Outcome: The proposed method improves on existing methods and provides accurate and computationally efficient uncertainty scores for sequence-level selective generation and claim-level fact-checking tasks.
Uncertainty Quantification for Large Language Models (2025.acl-tutorials)

Copied to clipboard

Challenge: Large language models (LLMs) produce hallucinations, which undermine user trust and reliability.
Approach: This tutorial offers the first systematic introduction to uncertainty quantification (UQ) for LLMs in text generation tasks.
Outcome: The proposed framework provides tools for communicating the reliability of a model answer.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations