Challenge: Existing methods to assess the correctness of RAG models fail to capture the model’s internal state during answer generation.
Approach: They propose a method to predict the correctness of RAG models by modeling the model’s uncertainty on quantified perturbations of input.
Outcome: Extensive experiments across multiple large language models show that the proposed approach quantifies RAG robustness by aligning predictions with ground truth with a MSE 0.002 while offering flexibility for diverse qualitative metrics.

Similar Papers

Why Uncertainty Estimation Methods Fall Short in RAG: An Axiomatic Analysis (2025.findings-acl)

Copied to clipboard

Challenge: Existing UE methods cannot reliably estimate the correctness of LLM responses in Retrieval-Augmented Generation (RAG) . Existing methods generate low uncertainty values without considering relevance of context to query .
Approach: They propose an axiomatic framework to identify deficiencies in existing UE methods and introduce five constraints that an effective UE method should meet after incorporating retrieved documents into the LLM’s prompt.
Outcome: The proposed framework satisfies all the axioms and improves correlation between uncertainty estimates and correctness.
QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for reducing LLM hallucinations rely on model-internal signals . Existing approaches rely only on model internal signals, resulting in unreliability .
Approach: They propose a method that shifts from subjective confidence to objective statistics . they leverage Infini-gram for millisecond-latency queries over 4 trillion tokens .
Outcome: The proposed method reduces hallucinations in large language models by reducing uncertainty in the model.
Retrieval-Augmented Generation with Estimation of Source Reliability (2025.emnlp-main)

Copied to clipboard

Challenge: Retrieval-Augmented Generation (RAG) is an effective approach to enhance the factual accuracy of large language models (LLMs).
Approach: They propose a multi-source RAG framework that estimates the reliability of sources and prioritizes highly reliable and relevant documents.
Outcome: The proposed framework outperforms baselines in scenarios with heterogeneous source reliability while scaling efficiently as the number of sources increases.
Faithfulness-Aware Uncertainty Quantification for Fact-Checking the Output of Retrieval-Augmented Generation (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to mitigating hallucinations conflate factuality with faithfulness to the retrieved evidence, incorrectly labeling factually correct statements as hallucinos . Existing methods to mitigate hallucinics rely on a lack of training data coverage, input ambiguity, and architectural constraints.
Approach: They propose a method for hallucination detection in Large Language Models enhanced with knowledge retrieval based on faithfulness to the retrieved context.
Outcome: The proposed method outperforms unsupervised UQ baselines, RAG-specific methods, and supervised classifiers across multiple tasks and LLMs.
Controlling Risk of Retrieval-augmented Generation: A Counterfactual Prompting Framework (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on retrieval-augmented generation (RAG) rarely address the issue of predictive uncertainty, i.e., how likely it is that a RAG model’s prediction is incorrect.
Approach: They propose a framework that induces RAG models to alter latent factors and analyzes the effect on their answers.
Outcome: The proposed framework identifies two critical factors affecting RAG models' confidence in their answers and analyzes the effect on their answers.
Classifying and Addressing the Diversity of Errors in Retrieval-Augmented Generation Systems (2026.eacl-long)

Copied to clipboard

Challenge: Existing work on RAG errors has not accounted for the complexity of real-world RAG systems and their failure modes.
Approach: They propose a taxonomy of error types that can occur in realistic RAG systems and an auto-evaluation method that can be used to track errors during development.
Outcome: The proposed method can be used in practice to track and address errors during development.
Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation Systems (2025.coling-main)

Copied to clipboard

Challenge: Retrieval-Augmented Generation (RAG) models address fairness concerns with respect to sensitive attributes such as gender, geographic location, and other demographic factors.
Approach: They propose a framework to evaluate fairness in RAG using scenario-based questions and analyzing disparities across demographic attributes.
Outcome: The proposed framework analyzes disparities across demographic attributes and identifies fairness issues in retrieval and generation stages.
Stable-RAG: Mitigating Retrieval-Permutation-Induced Hallucinations in Retrieval-Augmented Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing RAG methods focus on enhancing LLM robustness to low-quality retrieval, but neither address permutation sensitivity.
Approach: They propose a method that exploits permutation sensitivity to mitigate hallucinations in Large Language Models.
Outcome: The proposed model improves answer accuracy, reasoning consistency, and generalization across datasets, retrievers, and input lengths compared with strong baselines.
A Survey of RAG-Reasoning Systems in Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: a survey of RAG-based reasoning-based approaches shows that it is not effective for multi-step inferences.
Approach: They map how advanced reasoning optimizes each stage of RAG . they show how retrieved knowledge supply missing premises and expand context for complex inference .
Outcome: The proposed frameworks achieve state-of-the-art across knowledge-intensive benchmarks.
Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back Home (2025.acl-long)

Copied to clipboard

Challenge: Recent adaptive retrieval methods integrate LLMs’ intrinsic knowledge with external information appealing to LLM self-knowledge, but they often neglect efficiency evaluations and comparisons with uncertainty estimation techniques.
Approach: They propose to integrate LLMs’ intrinsic knowledge with external information appealing to LLM self-knowledge but neglect efficiency evaluations and comparisons with uncertainty estimation techniques.
Outcome: The proposed methods outperform complex pipelines in terms of efficiency and self-knowledge while maintaining comparable QA performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations