Challenge: Open domain question answering systems often rely on information retrieved from large collections of text to answer questions.
Approach: They evaluate and benchmark three powerful Large Language Models with a dataset . they find that 25% of unambiguous open domain questions can lead to conflicting contexts .
Outcome: The proposed model can't be used to answer questions with conflicting contexts . it can be fine tuned to provide richer information into the model's training .

Similar Papers

Adaptive Question Answering: Enhancing Language Model Proficiency for Addressing Knowledge Conflicts with Source Citations (2024.emnlp-main)

Copied to clipboard

Challenge: Existing work on citation generation has focused on unambiguous settings with single answers, failing to address the complexity of real-world scenarios.
Approach: They propose a task of QA with source citation in ambiguous settings where multiple valid answers exist, where multiple sources exist.
Outcome: The proposed framework generates multiple answers and cites their sources, allowing users to verify the factuality of each answer and make informed decisions.
DIVKNOWQA: Assessing the Reasoning Ability of LLMs via Open-Domain Question Answering over Knowledge Base and Text (2024.findings-naacl)

Copied to clipboard

Challenge: Retrievalaugmented LLMs have been used to ground LLM in external knowledge . a gap exists in the current landscape regarding the effectiveness of grounding LLM on heterogeneous knowledge sources.
Approach: They propose a model that uses symbolic language to generate symbolic queries . they use a dataset that is generated using predefined reasoning chains and human annotation .
Outcome: The proposed model outperforms previous approaches by a significant margin in QA tasks over text.
Blinded by Generated Contexts: How Language Models Merge Generated and Retrieved Contexts When Knowledge Conflicts? (2024.acl-long)

Copied to clipboard

Challenge: Recent advances in augmenting Large Language Models (LLMs) with auxiliary information have significantly revolutionized their efficacy in knowledge-intensive tasks.
Approach: They propose a systematic framework to identify whether LLMs’ responses are attributed to either generated or retrieved contexts.
Outcome: The proposed framework identifies whether LLMs’ responses are attributed to either generated or retrieved contexts.
DebateQA: Evaluating Question Answering on Debatable Knowledge (2026.findings-eacl)

Copied to clipboard

Challenge: Existing QA benchmarks that provide fixed answers to debatable questions are inadequate for evaluating their performance.
Approach: They propose to use a dataset of 2,941 debatable questions to assess their ability to provide comprehensive answers to inherently debatably asked questions.
Outcome: The proposed model performs well on 2,941 debatable questions accompanied by human-annotated partial answers that capture a variety of perspectives.
Open-Domain Question Answering (2020.acl-tutorials)

Copied to clipboard

Challenge: tutorial provides a comprehensive overview of cutting-edge research in open-domain question answering (QA)
Approach: tutorial provides a comprehensive overview of cutting-edge research in open-domain question answering . focus will shift to cutting- edge models proposed for open- domain QA .
Outcome: The tutorial will cover cutting-edge research in open-domain question answering (QA) it will cover two-stage retriever-reader approaches, dense retriever and end-to-end training, and retriever free methods .
Desiderata For The Context Use Of Question Answering Systems (2024.eacl-long)

Copied to clipboard

Challenge: Prior work has uncovered a set of common problems in state-of-the-art context-based question answering systems, such as a lack of attention to the context when it conflicts with a model’s parametric knowledge and a loss of consistency with their answers.
Approach: They propose to examine the desiderata for context-based question answering systems and then compare them to a set of prior work.
Outcome: The proposed models are based on 15 datasets and evaluated on 5 datasets.
AmbigQA: Answering Ambiguous Open-domain Questions (2020.emnlp-main)

Copied to clipboard

Challenge: Existing open-domain question answering systems assume questions have a single welldefined answer.
Approach: They propose an open-domain question answering task which involves finding every plausible answer and rewriting the question for each one to resolve the ambiguity.
Outcome: The proposed task is based on a dataset covering 14,042 open-domain questions . it shows that strong models benefit from weakly supervised learning .
QuAC: Question Answering in Context (D18-1)

Copied to clipboard

Challenge: a dataset for Question Answering in Context contains 14K information-seeking QA dialogs . questions are often more open-ended, unanswerable, or only meaningful within the dialog context .
Approach: They propose a dataset for Question Answering in Context that contains 14K dialogs . they use a student to ask questions about a Wikipedia section and a teacher to answer them .
Outcome: The proposed dataset underperforms humans in a number of reference models . the dataset contains 14K information-seeking dialogs over sections from Wikipedia .
Conflicting Needles in a Haystack: How LLMs behave when faced with contradictory information (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive capabilities in retrieving and analyzing complex information, but their reliability in conflicting contexts remains poorly understood.
Approach: They propose an adversarial extension of the Needle-in-a-Haystack framework in which three mutually exclusive “needles” are embedded within long documents.
Outcome: The proposed framework highlights critical limitations in the robustness of current LLMs—including commercial systems—to contradiction.
Benchmarking LLM’s Capability in Reasoning over Conflicting Web References (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) integrated with retrieval-augmented generation (RAG) are a dominant framework for building intelligent assistants.
Approach: They propose a benchmark to evaluate LLMs' reasoning capability over real-world conflicting documents retrieved from the web.
Outcome: The proposed benchmark evaluates LLMs' reasoning capability over real-world conflicting documents retrieved from the web.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations