Challenge: Retrieval-Augmented Generation (RAG) systems are limited in their evaluation due to the intricate interplay between retrieval and generation components.
Approach: They propose a Question Answering Question Answerer dataset specifically designed for RAG evaluation that integrates external, non-parametric knowledge retrieved by a retrieval pool of 37,800 entries.
Outcome: The proposed dataset consists of 7,560 curated instances mapped to a retrieval pool of 37,800 entries, enabling an efficient evaluation of both retrieval and generation tasks.

Similar Papers

Benchmarking Retrieval-Augmented Generation for Medicine (2024.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have state-of-the-art performance on a wide range of medical question answering tasks, but they still face challenges with hallucinations and outdated knowledge.
Approach: They propose a benchmark to evaluate medical RAG systems using large-scale experiments with over 1.8 trillion prompt tokens.
Outcome: The proposed benchmark improves accuracy of six different LLMs by up to 18% over chain-of-thought prompting.
MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems (2025.naacl-long)

Copied to clipboard

Challenge: Traditional retrieval-augmented generation benchmarks use heuristics as the ground truth for evaluation, but require an expensive large language model (LLM) as a judge for a reliable evaluation.
Approach: They propose to use large language models as a judge for retrieval-augmented generation benchmarks . they use heuristic metrics as input and a large language model as heuriistic input .
Outcome: The proposed method couples heuristic features with large language models as judge for evaluation.
BERGEN: A Benchmarking Library for Retrieval-Augmented Generation (2024.findings-emnlp)

Copied to clipboard

Challenge: Retrieval-Augmented Generation allows to enhance Large Language Models with external knowledge.
Approach: They propose a library that allows to benchmark and standardize RAG experiments.
Outcome: The proposed library is an end-to-end library for reproducible research standardizing RAG experiments.
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmented Generation (2025.findings-naacl)

Copied to clipboard

Challenge: Existing research focuses on single-turn RAG, leaving a gap in addressing multi-turn conversations . a new benchmark is designed to assess RAG systems in realistic multi-turned conversations based on Wikipedia .
Approach: They propose a large-scale benchmark to assess RAG systems in multi-turn contexts . CORAL includes diverse information-seeking conversations automatically derived from Wikipedia . authors propose unified framework to standardize various conversational RAG methods .
Outcome: The proposed framework supports three core tasks of conversational RAG: passage retrieval, response generation, and citation labeling.
RAGAs: Automated Evaluation of Retrieval Augmented Generation (2024.eacl-demo)

Copied to clipboard

Challenge: RAGAs are a framework for reference-free evaluation of Retrieval Augmented Generation (RAG) pipelines.
Approach: They propose a framework for reference-free evaluation of Retrieval Augmented Generation pipelines.
Outcome: RAGAs can be used to evaluate RAG pipelines without human annotations . the framework can be useful for faster evaluation cycles given the fast adoption of LLMs based on human annotation.
RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluation metrics for RAG systems are lacking due to high costs of data construction and lack of factual accuracy.
Approach: They propose a framework to evaluate RAG systems in specialized scenarios . they propose three new metrics to evaluate LLM-generated responses .
Outcome: The proposed framework outperforms zero-shot and one-shot methods in terms of clarity, safety, conformity, and richness of generated samples.
RAG in the Wild: On the (In)effectiveness of LLMs with Mixture-of-Knowledge Retrieval Augmentation (2026.findings-acl)

Copied to clipboard

Challenge: Retrieval-augmented generation (RAG) enhances large language models by integrating external knowledge retrieved at inference time.
Approach: They evaluate RAG systems using MassiveDS, a large-scale datastore with mixture of knowledge.
Outcome: The proposed approach improves performance on knowledge-intensive NLP tasks.
Enhancing Retrieval-Augmented Generation: A Study of Best Practices (2025.coling-main)

Copied to clipboard

Challenge: Retrieval-augmented generation systems have shown remarkable advancements by integrating retrieval mechanisms into language models, enhancing their ability to produce more accurate and contextually relevant responses.
Approach: They propose to integrate query expansion, various novel retrieval strategies, and a Contrastive In-Context Learning RAG to improve response quality.
Outcome: The proposed RAGs incorporate query expansion, various novel retrieval strategies, and a novel Contrastive In-Context Learning RAG.
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems (2024.naacl-long)

Copied to clipboard

Challenge: Evaluating retrieval-augmented generation systems relies on hand annotations for input queries, passages to retrieve, and responses to generate.
Approach: They propose an automated evaluation framework for retrieval-augmented generation (RAG) ARES fine tunes lightweight LLM judges on synthetically generated queries and answers .
Outcome: The proposed framework evaluates RAG systems using only human annotations . it can be used to improve system understanding and create targeted solutions .
Know Your RAG: Dataset Taxonomy and Generation Strategies for Evaluating RAG Systems (2025.coling-industry)

Copied to clipboard

Challenge: Retrieval Augmented Generation (RAG) systems are widespread in the industry.
Approach: They propose to use Q&A datasets to assess retrieval performance and label-targeted data generation to refine RAG datasets.
Outcome: The proposed system can generate Q&A datasets with fine-tuned small LLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations