Challenge: Several benchmarks have been released to evaluate model performance on multi-turn retrieval augment generation tasks.
Approach: They propose to benchmark 666 conversations with over 2,800 conversation turns across 6 domains and a corpora that focuses on unanswerable questions and later conversation turns.
Outcome: The proposed benchmarks show that retrieval and generation models struggle on conversations with UNanswerable, UNderspecified, and NONstandalone questions and UNclear responses.

Similar Papers

CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmented Generation (2025.findings-naacl)

Copied to clipboard

Challenge: Existing research focuses on single-turn RAG, leaving a gap in addressing multi-turn conversations . a new benchmark is designed to assess RAG systems in realistic multi-turned conversations based on Wikipedia .
Approach: They propose a large-scale benchmark to assess RAG systems in multi-turn contexts . CORAL includes diverse information-seeking conversations automatically derived from Wikipedia . authors propose unified framework to standardize various conversational RAG methods .
Outcome: The proposed framework supports three core tasks of conversational RAG: passage retrieval, response generation, and citation labeling.
MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have greatly enhanced dialogue systems, but evaluation of their capabilities remains a challenge.
Approach: They propose a model to evaluate the fine-grained abilities of Large Language Models in multi-turn dialogues.
Outcome: The proposed model evaluates 21 popular chatbots based on MT-Bench-101 . it includes 3 overarching abilities and 13 distinct tasks within multi-turn dialogue scenarios.
MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on single-turn evaluations, overlooking the models’ capabilities in multi-turn interactions.
Approach: They propose a benchmark to evaluate the multi-turn conversational abilities of large language models (LLMs) by analyzing human-LLM conversations and constructing multi-turned queries for each category using GPT-4.
Outcome: The proposed model outperforms open-source models in multi-turn tasks while retaining and recalling historical information.
RAD-Bench: Evaluating Large Language Models’ Capabilities in Retrieval Augmented Dialogues (2025.naacl-industry)

Copied to clipboard

Challenge: Existing benchmarks assess LLMs' chat abilities in multi-turn dialogues or their use of retrieval for augmented responses in limited tasks such as knowledge QA or numeric reasoning.
Approach: They propose a benchmark to evaluate LLMs' capabilities in multi-turn dialogues following retrievals.
Outcome: The proposed benchmark evaluates LLMs' ability to perform in multi-turn dialogues following retrievals over 6 representative scenarios.
Comprehensive Comparison of RAG Methods Across Multi-Domain Conversational QA (2026.eacl-srw)

Copied to clipboard

Challenge: Existing studies evaluate RAG methods in isolation and focus on single-turn settings.
Approach: They compare retrieval-augmented generation methods for multi-turn conversational QA with those that use dialogue history and coreference to ground large language models.
Outcome: The proposed methods outperform vanilla RAG and advanced methods fail to yield gains and can even degrade performance below the No-RAG baseline.
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have shown promising results in complex reasoning tasks.
Approach: They propose to use a multi-turn reasoning evaluation framework to cover multi-turn interactions with the environments of large language models.
Outcome: The proposed framework covers diverse reasoning capabilities, fine-grained difficulty granularity, and necessitates multi-turn interactions with the environments.
XRAG: Cross-lingual Retrieval-Augmented Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: XRAG evaluates the generation abilities of LLMs in cross-lingual RAG settings where the user language does not match retrieval results.
Approach: They propose a benchmark to evaluate the generation abilities of LLMs in cross-lingual RAG settings where the user language does not match retrieval results.
Outcome: XRAG is a benchmark designed to evaluate the generation abilities of LLMs in cross-lingual RAG settings where the user language does not match retrieval results.
MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks (2026.acl-long)

Copied to clipboard

Challenge: Existing conversational retrieval benchmarks suffer from costly, sparse human annotation or rigid, unnatural automated heuristics.
Approach: They propose a framework for auditing, synthesizing, and benchmarking conversational retrieval.
Outcome: The proposed framework is based on three LLM-based auditors and a multi-agent system . it mimics production-style challenges (hard topic switching, verbosity) and offers superior discriminative power.
CoRe-MMRAG: Cross-Source Knowledge Reconciliation for Multimodal RAG (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in Multimodal Large Language Models have significantly improved reasoning and generation tasks by leveraging joint vision-language representations.
Approach: They propose a framework that reconciles inconsistencies across knowledge sources . they use a four-stage pipeline to generate an internal response from parametric knowledge .
Outcome: Experiments on KB-VQA show that CoRe-MMRAG achieves performance gains of 5.6% and 9.3% over baseline methods.
MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text (2022.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models store a massive amount of world knowledge implicitly in their parameters, but large models often fail to encode information about rare entities and events.
Approach: They propose a retrieval-augmented model which accesses an external non-parametric memory to augment language generation.
Outcome: The proposed model outperforms existing models by 10-20% absolute on two datasets and under distractor and full-wiki settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations