Papers with Question Answering
Copied to clipboard
| Challenge: | Mapping and navigation services struggle to handle natural language geospatial queries. |
| Approach: | They introduce an extensible open-source framework that streamlines the creation of reproducible, traceable map-based QA datasets. |
| Outcome: | a new open-source framework streamlines the creation of reproducible, traceable map-based QA datasets. |
Copied to clipboard
| Challenge: | Existing benchmarks for large language models focus on intradocument dependencies or dependencies between a small number of documents. |
| Approach: | They propose to use a dataset of fan-out question-answer pairs and human-annotated decompositions with English Wikipedia as the knowledge base to evaluate models' reasoning. |
| Outcome: | The proposed dataset shows that models still have room to improve reasoning over inter-document dependencies in a long context. |
Copied to clipboard
| Challenge: | Existing tools for Question Answering (QA) have challenges that limit their use in practice. |
| Approach: | They propose a library that integrates with existing infrastructure and offers helpful defaults for QA subtasks. |
| Outcome: | NeuralQA integrates well with existing infrastructure and offers helpful defaults for QA subtasks. |
Copied to clipboard
| Challenge: | ARC-Easy, ARC Challenge, and OpenBookQA use Wikipedia to augment training data . performance degrades when additional instances exhibit higher difficulty than original training data. |
| Approach: | They propose two methods for exploiting external knowledge for QA in science . they enrich the original corpus with relevant text snippets from an open-domain resource . the second method simply increases the amount of training data by appending additional in-domain instances. |
| Outcome: | The proposed methods achieve gains in accuracy of 8.1%, 13.0%, and 12.8% on science QA tasks. |
Copied to clipboard
| Challenge: | Talk to Papers aims to improve the current experience of academic search by using open-domain question answering (QA) techniques. |
| Approach: | They propose to use open-domain question answering techniques to improve the current experience of academic search by combining natural language queries with machine reading at scale. |
| Outcome: | The proposed tool improves on existing search engines and provides a collaborative data collection tool to curate the first natural language processing research QA dataset. |
Copied to clipboard
| Challenge: | Existing studies to solve QA tasks in an integrated manner are not available in other languages because of the lack of QA datasets. |
| Approach: | They build a Japanese version of Natural Questions using natural questions from query logs of a search engine and crowdsource it using crowdsourcing. |
| Outcome: | The proposed datasets are based on natural questions from Japanese search engines and crowdsourced. |
Copied to clipboard
| Challenge: | Standardized tests have been proposed as replacements to the Turing test as a driver for progress in AI. |
| Approach: | et al. propose standardized tests as replacements to the Turing test as a driver for progress in AI. |
| Outcome: | a series of standardized tests have been proposed as replacements to the Turing test . the tutorial categorizes open domain and closed domain tests into two categories . open domain tests require the system to have significant domain knowledge and reasoning capabilities. |
Copied to clipboard
| Challenge: | Question Answering (QA) is a major area of research in Natural Language Processing (NLP) |
| Approach: | They propose a one-stop and open-source QA repository for question answering . it supports core QA functionalities like retrieval and reading comprehension . they say it will facilitate easy replication of state-of-the-art (SOTA) QA methods . |
| Outcome: | The proposed framework enables easy replication of state-of-the-art (SOTA) QA methods. |
Copied to clipboard
| Challenge: | despite considerable progress, most machine reading comprehension tasks lack sufficient training data to fully exploit powerful deep neural network models. |
| Approach: | They propose to use QA data to generate more training data for machine reading comprehension tasks by crowdsourcing . they first collect a large-scale multiple-choice QA dataset for Chinese, ExamQA, and then use incomplete, yet relevant snippets returned by a web search engine as the context for each QA instance. |
| Outcome: | The proposed model improves a Chinese MRC task with +5.1% accuracy and +3.8% exact match. |
Copied to clipboard
| Challenge: | Current Large Language Models (LLMs) benchmarks are often based on open-ended or close-ended QA evaluations, avoiding the requirement of human labor. |
| Approach: | They propose a multi-axis suite for healthcare LLM evaluation, exploring correlations between open and close benchmarks and metrics. |
| Outcome: | The proposed framework explores correlations between open and close benchmarks and metrics in the healthcare domain, with blind spots and overlaps in existing methodologies. |
Copied to clipboard
| Challenge: | Using a benchmark, large language models can be evaluated on Moroccan legal MCQs . despite their ability to comprehend and process Arabic, the language is still a challenge . |
| Approach: | They propose a benchmark for assessing LLMs on Moroccan legal MCQs . they use Arabic-based questions enriched with Moroccan idioms to assess their accuracy . |
| Outcome: | The proposed benchmark covers 1,776 expert-verified questions in Arabic enriched with Moroccan idioms . it measures accuracy, precision-penalized F1-like score, and calibration errors . |
Copied to clipboard
| Challenge: | ALCQA addresses the semantic and structural gap between natural language and action sequences . a priori, the semantics of the question and action are not well understood . |
| Approach: | They propose an alignment-enhanced complex question answering framework which aligns questions and actions into sequences. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on a CQA and WQSP dataset. |
Copied to clipboard
| Challenge: | Abstractive summarization systems often generate summaries with factual errors . many approaches to detect these errors have been proposed, but this capability has not been evaluated in past research . |
| Approach: | They propose to use question answering-based factuality metrics to detect errors in summaries . they find that QA-based frameworks fail to correctly identify error spans in generated summary . |
| Outcome: | The proposed methods outperform trivial exact match baselines in localizing errors in summaries. |
Copied to clipboard
| Challenge: | Existing models for multihop reasoning are limited in their performance . multi-hop reasoning requires the ability to gather information from multiple passages . |
| Approach: | They propose a method that provides the full reasoning chain of multiple passages instead of just one final passage where the answer appears. |
| Outcome: | The proposed model improves on existing models by providing the full reasoning chain of multiple passages instead of just one final passage where the answer appears. |
Copied to clipboard
| Challenge: | Existing approaches to document QA use a pre-retrieval step to retrieve the relevant context from documents, but this is incongruous with the user's mental model of the document. |
| Approach: | They propose an approach called PDFTriage that enables models to retrieve the context based on either structure or content. |
| Outcome: | The proposed approach can retrieve context based on structure or content across several classes of questions where existing retrieval-augmented LLMs fail. |
Copied to clipboard
| Challenge: | Understanding natural language questions entails the ability to break down a question into the requisite steps for computing its answer. |
| Approach: | They introduce a Question Decomposition Meaning Representation (QDMR) for questions . they demonstrate that QDMRs can be annotated at scale using a hotpotQA dataset . |
| Outcome: | The proposed model outperforms several natural baselines in the open-domain question answering hotpotQA dataset and can be deterministically converted to a pseudo-SQL formal language. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have significant potential for facilitating intelligent end-user applications in healthcare, but hallucinations remain an inherent problem with LLMs. |
| Approach: | They propose a pipeline to retrieve and medicalally-augmented-generation with knowledge reduction using cross-encoder re-ranking strategies to reduce the knowledge base. |
| Outcome: | The proposed pipeline reduces the knowledge base and improves inference time by 47%. |
Copied to clipboard
| Challenge: | a new open-domain question answering system integrates best practices from IR with a BERT-based reader to identify answers from a large corpus of Wikipedia articles. |
| Approach: | They propose an end-to-end question answering system that integrates BERT with an IR reader. |
| Outcome: | The proposed system improves on a standard benchmark test collection. |
Copied to clipboard
| Challenge: | DOCMASTER is a platform for annotating PDF documents, model training, and inference, tailored to document question-answering. |
| Approach: | They propose to integrate layout information into a unified platform for annotating PDF documents, model training, and inference tailored to document question-answering. |
| Outcome: | The proposed platform is designed for annotating PDF documents, model training, and inference, tailored to document question-answering. |
Copied to clipboard
| Challenge: | Reading comprehension models are dominated by recurrent neural networks (RNNs) as documents become longer and questions become complex, sequential reading becomes a significant bottleneck. |
| Approach: | They propose a reading comprehension framework that uses document trees to model an agent that interleaves quick navigation with more expensive answer extraction. |
| Outcome: | The proposed model improves question answering performance compared to existing models and has a strong information-retrieval baseline. |
Copied to clipboard
| Challenge: | Multi-hop question answering (QA) requires an information retrieval system that can find multiple supporting evidence needed to answer the question. |
| Approach: | They propose a technique that uses information of entities present in the initial retrieved evidence to learn to ‘hop’ onto other relevant evidence. |
| Outcome: | The proposed method boosts retrieval performance on a multi-hop question answering dataset with 5 million Wikipedia paragraphs and a model without training increases its performance by 10.59 F1. |
Copied to clipboard
| Challenge: | scalability of attribute-value extraction (AVE) task is key for a large number of products . a question-answering (QA)-based approach is better for AVE, but requires a larger number of classes to be scalable. |
| Approach: | They propose a question-answering-based approach that additionally inputs the target attribute as a query to extract its values. |
| Outcome: | The proposed approach outperforms a classical approach on real-word e-commerce datasets in accuracy and speed. |
Copied to clipboard
| Challenge: | Existing n-gram based QA metrics have a number of drawbacks and are not suitable for all extractive tasks. |
| Approach: | They propose to use BERTScore to evaluate translation for question answering (QA) they also explore whether existing n-gram based metrics are suitable for generative QA . |
| Outcome: | The proposed BERTScore metric fails to provide stronger correlation with human judgements . |
Copied to clipboard
| Challenge: | Recent advances in machine reading have inspired researchers to combine Information Retrieval with machine reading to tackle open-domain QA. |
| Approach: | They propose two neural network rankers that assign scores to different passages based on their likelihood of containing the answer to a given question. |
| Outcome: | The proposed models achieve human level performance in open-domain QA compared to reading comprehension-style QA because it is difficult to retrieve the pieces of paragraphs that contain the answer to the question. |
Copied to clipboard
| Challenge: | a gap remains in reasoning ability compared to a human, and performance tends to degrade when models are exposed to less-constrained tasks. |
| Approach: | They conduct extensive qualitative and quantitative analyses on the results of four models across four datasets . they relate common errors to model capabilities and discuss a way forward . |
| Outcome: | The proposed model performance is based on the results of four models across four datasets. |
Copied to clipboard
| Challenge: | Existing question answering datasets are imperfect tests that do not expose model limitations. |
| Approach: | They develop an adversarial writing setting where humans interact with trained models and try to break them. |
| Outcome: | The proposed model-driven annotation process systematically stumps automated question answering systems. |
Copied to clipboard
| Challenge: | Existing models for question answering are limited in the availability of labeled data. |
| Approach: | They propose a hierarchical conditional variational autoencoder for generating QA pairs given unstructured texts as contexts while maximizing mutual information between generated QA pair to ensure consistency. |
| Outcome: | The proposed framework achieves impressive performance gains over baseline models on both tasks, using only a fraction of data for training. |
Copied to clipboard
| Challenge: | Existing approaches to predict product-related questions fail for new or unpopular products . product-specific question answering is a popular service provided by many e-commerce websites . |
| Approach: | They propose a framework for predicting the answer to product-related questions based on the answers of similar products. |
| Outcome: | The proposed model outperforms baselines on some segments of product-related questions. |
Copied to clipboard
| Challenge: | Community Question Answering is a research area that benefits from deep linguistic analysis . previous cQA challenges have shown that neural approaches are not enough to deliver state-of-the-art results . |
| Approach: | They propose a framework to distribute computation of cQA tasks over computer clusters . community question answering is a research area that benefits from deep linguistic analysis . |
| Outcome: | The proposed framework scales to large datasets and delivers fast processing. |
Copied to clipboard
| Challenge: | ClinicalTrialsHub consolidates clinical trial data from ClinicalTrial.gov and augments it by extracting and structuring trial-relevant information from PubMed. |
| Approach: | They propose a search-focused platform that consolidates PubMed data and extracts structured trial information. |
| Outcome: | ClinicalTrialsHub increases access to structured clinical trial data by 83.8% compared to ClinicalTrial.gov alone. |
Copied to clipboard
| Challenge: | Adapting models to new domain without finetuning is a challenging problem in deep learning. |
| Approach: | They propose an adversarial training framework for domain generalization in Question Answering task using a conventional QA model and a discriminator. |
| Outcome: | The proposed model outperforms the baseline model on Question Answering (QA) task. |
Copied to clipboard
| Challenge: | ComQA dataset captures question phenomena and the diverse ways in which they are formulated. |
| Approach: | They propose a large dataset of real user questions that captures question phenomena and the diverse ways in which they are formulated. |
| Outcome: | The proposed dataset can be a driver of future research on factoid question answering (QA). |
Copied to clipboard
| Challenge: | Existing approaches to open-domain question answering struggle to retrieve indirectly related evidence when no direct evidence is provided. |
| Approach: | They propose a retriever-reader model that learns to attend on essential terms during the question answering process. |
| Outcome: | The proposed model achieves the state-of-the-art on multiple open-domain QA datasets and achieves a 'reader-reader' level. |
Copied to clipboard
| Challenge: | Existing multihop reasoning benchmarks are largely solvable via shortcuts . a bottom–up approach allows us to create a multihop QA dataset that requires proper multihop thinking. |
| Approach: | They propose a bottom–up approach that selects composable pairs of single-hop questions that are connected and adds stringent filters to the construction process. |
| Outcome: | The proposed approach creates a multihop question answering dataset with 25K 2–4 hop questions. |
Copied to clipboard
| Challenge: | Existing abstractive question-answering datasets in Vietnamese are lacking . |
| Approach: | They propose to introduce a Vietnamese abstractive question-answering corpus to address this gap . they propose to use Vietnamese abstractives to generate answers to questions . |
| Outcome: | The proposed dataset examines the capability of large language models in the Vietnamese medical domain, including reasoning, memorizing and awareness of essential information. |
Copied to clipboard
| Challenge: | Existing QA datasets rarely distinguish fine-grained reading skills, such as the understanding of varying narrative elements. |
| Approach: | They propose to use FairytaleQA to generate 10,580 questions based on 278 children-friendly stories to assess model's fine-grained learning skills. |
| Outcome: | The proposed dataset consists of 10,580 questions derived from 278 children-friendly stories, covering seven types of narrative elements or relations. |
Copied to clipboard
| Challenge: | Large language models have achieved high performance on various natural language benchmarks, but the explainability of their output remains elusive. |
| Approach: | They propose an architecture called iterative retrieval-generation reasoner that generates an entailment tree that explains a given hypothesis by using premises from C. |
| Outcome: | The proposed model outperforms existing benchmarks on premise retrieval and entailment tree generation with around 300% gain in overall correctness. |
Copied to clipboard
| Challenge: | Prior work has focused on the case in which the question is clear and the answer is unambiguous but possibly unknown. |
| Approach: | They propose to use a sampled set of questions to calibrate answers to ambiguous questions with varying model scales. |
| Outcome: | The results show that sampling-based confidence scores help calibrate answers to relatively unambiguous questions, with more dramatic improvements on ambiguous ones. |
Copied to clipboard
| Challenge: | a growing body of research suggests that social descriptors can influence LLM-generated clinical recommendations. |
| Approach: | They examine whether social descriptors of a patient distort uncertainty signals and model accuracy. |
| Outcome: | The presence of social identity cues affects the reliability of confidence signals, the authors show . incorporating sociodemographic attributes alters outputs in clinical trial matching and QA . |
Copied to clipboard
| Challenge: | Existing models are far from perfect when assessed at the level of clusters of semantically connected probes, such as all hypernym questions about a single concept. |
| Approach: | They propose a method for automatically building probe datasets from expert knowledge sources, allowing systematic control and a comprehensive evaluation. |
| Outcome: | The proposed model is predisposed to recognize certain types of structural linguistic knowledge, but performance degrades even with a slight increase in the number of “hops” in the underlying taxonomic hierarchy. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are trained on vast corpora that contain substantial knowledge but their outputs often contain confidently stated inaccuracies. |
| Approach: | They propose to encode truthfulness as a distinct linear feature, termed the "truth direction", which can classify truthfulness reliably. |
| Outcome: | The proposed model can generalize to logical transformations, question-answering tasks, in-context learning, and external knowledge sources. |
Copied to clipboard
| Challenge: | Existing research on niche answer types, mainly short responses and, in a few cases, long responses, has failed to adequately address the answer diversity of questions. |
| Approach: | They propose to use Google's autocomplete feature to collect questions from a large-scale dataset with a variety of answer types to facilitate further research on improving QA with diverse response types. |
| Outcome: | The proposed model produces naturalistic questions that are short and expressed using simple language. |
Copied to clipboard
| Challenge: | Existing work on automating financial numerical reasoning focuses on unrealistically specific document snippets, failing to reflect the broader and more realistic scenarios faced by analysts. |
| Approach: | They propose a long-document financial QA task that augments 7,437 questions from existing FinQA dataset with full-document context, extending the average context length from under 700 words in FinQA to 123k words in DocFinQA. |
| Outcome: | The proposed task extends the average context length from under 700 words in FinQA to 123k words in DocFinQA. |
Copied to clipboard
| Challenge: | Existing QA systems that answer factual questions with short answers are rare in practice. |
| Approach: | They propose a proposed two-layered taxonomy technique for semantic question matching . they augment state-of-the-art deep learning models with question classes from a deep learning based question classifier . |
| Outcome: | The proposed technique achieves state-of-the-art on an open-domain dataset. |
Copied to clipboard
| Challenge: | Existing models for text-based multiple choice question answering are based on a text. |
| Approach: | They propose a Convolutional Neural Network (CNN) model for text-based multiple choice question answering where questions are based on a particular article. |
| Outcome: | The proposed model outperforms several baseline models on the SciQ and TQA datasets. |
Copied to clipboard
| Challenge: | a dataset of 40k information-seeking questions across seven languages is used to answer multilingual question answering tasks. |
| Approach: | They propose a task framework that allows questions from one language to be answered via answer content from another language. |
| Outcome: | The proposed framework can be used to answer questions from one language to another . the dataset was built on 40K questions across 7 languages, but could not find same-language answers . |
Copied to clipboard
| Challenge: | Large language models excel at financial reasoning but their deployment for enterprise use cases remains costly and often constrained by latency, privacy, and regulatory requirements. |
| Approach: | They propose a pipeline that extracts and selects relevant content from unstructured financial documents and generates QA pairs from the selected content for SLM fine-tuning. |
| Outcome: | The proposed model outperforms models trained on previous manual models and achieves competitive in-distribution performance. |
Copied to clipboard
| Challenge: | Existing QA models rely on learning interaction between document and question . current models require explicit attention to the document before or as it reads it . |
| Approach: | They propose a modular question answering task that enforces complete independence of the document encoder from the question encoder. |
| Outcome: | The proposed model achieves reasonable accuracy but significantly underperforms unconstrained QA models. |
Copied to clipboard
| Challenge: | Recent work has combined open-domain question answering with machine comprehension models to find answers in a large knowledge source. |
| Approach: | They propose a machine comprehension model that ranks paragraphs of retrieved documents for a higher answer recall with less noise. |
| Outcome: | The proposed model improves on four open-domain QA datasets by 7.8% on average. |
Copied to clipboard
| Challenge: | Question answering is one of the most common tasks in natural language processing . open-domain questions cover a wide range of topics and do not necessarily come in form of an actual question. |
| Approach: | They describe a Russian question-like question set collected from the Russian analogue of Jeopardy! They observe its linguistic features and the related QA-task. |
| Outcome: | The proposed data set includes 379,284 quiz-like questions with 29,375 from the Russian analogue of Jeopardy! |
Copied to clipboard
| Challenge: | Existing methods for POI oriented question answering lack ability to handle important POI related information. |
| Approach: | They propose a deep learning framework integrated with joint inference to capture tag semantic and geographic correlation between question and POIs. |
| Outcome: | The proposed model captures both tag semantic and geographic correlation between question and POIs. |
Copied to clipboard
| Challenge: | Visual Question Answering (VQA) is increasingly used in diverse applications where models must provide accurate answers and explanations that humans can easily understand and verify. |
| Approach: | They propose a unified prototypical framework that learns question-aware prototypes that serve as reasoning anchors and applies spatially constrained matching to ensure that the selected evidence is coherent and semantically relevant. |
| Outcome: | The proposed framework yields faithful, fine-grained explanations while maintaining competitive accuracy. |
Copied to clipboard
| Challenge: | Despite remarkable progress made in natural language processing, even the state-of-the-art systems often make incorrect predictions. |
| Approach: | They propose to use selective prediction to enable models to abstain from answering when their predictions are likely to be incorrect. |
| Outcome: | The proposed method improves performance on 11 QA datasets and in- and out-of-domain settings. |
Copied to clipboard
| Challenge: | Neural generation models often struggle to identify which content units are salient. |
| Approach: | They propose a new conceptualization of text plans as a sequence of question-answer pairs . they propose QA blueprints as QA proxy for content selection and planning . |
| Outcome: | The proposed model improves existing datasets with QA blueprints as proxy for content selection and planning. |
Copied to clipboard
| Challenge: | Existing methods for question answering and question generation are hard to obtain in many domains. |
| Approach: | They propose a method for jointly learning to ask and answer questions . they leverage unlabeled text along with labeled question answer pairs for learning . |
| Outcome: | The proposed method improves on four benchmark datasets on question answering and question generation tasks. |
Copied to clipboard
| Challenge: | Evaluating Question Answering systems in low-resource Indic languages remains challenging due to the scarcity of annotated data and the lack of reliable evaluation metrics. |
| Approach: | They propose a language-based multi-aspect evaluation framework for question answering systems . the framework integrates semantic similarity, factual completeness, numerical accuracy and contextual relevance . |
| Outcome: | The proposed metric is evaluated across eight Indic-language QA tasks using multiple LLMs . Across all settings, it shows stronger agreement with human evaluation . |
Copied to clipboard
| Challenge: | Current Visual Question Answering (VQA) models are trained on labelled data that may be insufficient to learn complex knowledge representations. |
| Approach: | They propose a method to integrate external knowledge into a visual pre-trained model by integrating facts extracted from a knowledge base. |
| Outcome: | The proposed method outperforms baseline models on the KVQA dataset benchmark by 19% and shows that it is weaker than previous models. |
Copied to clipboard
| Challenge: | Existing product question answering models do not provide labelled data for the task and description information for products is very lengthy. |
| Approach: | They propose a distant supervision-based NLI model to prepare training data without manual efforts. |
| Outcome: | The proposed model outperforms standard multi-task fine-tuning and improves 6% in human evaluation over baselines. |
Copied to clipboard
| Challenge: | Large language models (LLMs) and their applications in low-resource languages are limited due to lack of training data and benchmarking datasets. |
| Approach: | They propose a question-response system for Vietnamese that uses LLMs . they propose to open-source the model and train it on benchmark datasets based on Vietnamese data . |
| Outcome: | The proposed question answering system for Vietnamese is open-source and performant . it can learn and capture human-like text, but there is a gap in evaluations for Vietnamese . |
Copied to clipboard
| Challenge: | Earth Virtual Expert (EVE) is the first open-source, end-to-end initiative for developing and deploying domain-specialized LLMs for Earth Intelligence. |
| Approach: | They introduce Earth Virtual Expert, an open-source initiative for developing and deploying domain-specialized LLMs for Earth Intelligence. |
| Outcome: | The proposed model outperforms existing models on Earth Observation and Earth Sciences benchmarks while maintaining general capabilities. |
Copied to clipboard
| Challenge: | Existing systems that use contradiction to determine if a question is supported by background contexts do better than those that use entailment. |
| Approach: | They propose a method that incorporates contradiction in natural language inference (NLI) they propose to reformulate answers from QA systems as hypotheses and then select the best one based on the results. |
| Outcome: | The proposed method improves on multiple choice and extractive QA in two settings. |
Copied to clipboard
| Challenge: | LLMs are widely used for information seeking, but their generated responses often suffer from hallucinations, hindering their widespread adoption in high stakes domains such as law. |
| Approach: | They propose to attribute legal question answering to an actual source to improve factuality and verifiability of the answer. |
| Outcome: | The proposed framework improves the factuality and verifiability of legal question answering by combining a dataset from ECHR case law guides with an LLM-based filtering pipeline. |
Copied to clipboard
| Challenge: | Diagram question answering (Diagram QA) requires reasoning-level attribution that links each question-answer pair to all visual regions needed to derive the answer. |
| Approach: | They propose a diagram question-answer review framework that decouples interface logic from dataset-specific JSON structures through an internal meta-schema and dataset adapters. |
| Outcome: | The proposed framework achieves 85.39% precision and 75.30% recall across six diagram QA datasets. |
Copied to clipboard
| Challenge: | Question Answering (QA) has primarily focused on knowledge bases or free text as a source of knowledge. |
| Approach: | They propose a task of multi-relational QA over personal narrative using text worlds . they generate and release a lightweight Python-based framework for easily generating additional worlds and narrative . |
| Outcome: | The proposed framework combines elements of structured QA over knowledge bases and unstructured QA . it generates and analyzes five diverse datasets with dynamic narrative . the framework is lightweight and easy to use . |
Copied to clipboard
| Challenge: | Existing approaches to predicting examinee proficiency from short-answer questions (SAQs) use of labeled data to train on is difficult, and requires expensive expert-rated data. |
| Approach: | They propose a method to predict examinee proficiency from short-answer questions . previous approaches train on manually labeled data to predict human-ratings assigned to SAQs . |
| Outcome: | The proposed model examines examinee proficiency directly and does not require manual training on labeled data. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) and Retrieval-augmented Generation (RAG) systems show promise, but their performance on cross-document MEQA remains underexplored due to the lack of tailored benchmarks. |
| Approach: | They propose a scalable multi-document, multi-entity benchmark to evaluate LLMs' capacity to retrieve, consolidate, and reason over scattered and dense information. |
| Outcome: | The proposed benchmarks show that even advanced models achieve only 59% accuracy on MEBench. |
Copied to clipboard
| Challenge: | Existing datasets for Yes/No QA are lacking information needed to answer a Yes/Non question. |
| Approach: | They extend the Yes/No QA task by adding questions with an IDK answer to a BoolQ dataset and create out-of-domain test sets for the task. |
| Outcome: | The proposed dataset includes paragraphs together with naturally occurring questions whose answer is either "Yes" or "No". |
Copied to clipboard
| Challenge: | Existing language models that answer recipes better than humans can mitigate risks to users. |
| Approach: | They propose to specialize the analysis to more concrete applications and their plausible users. |
| Outcome: | The proposed model answers recipes as well or better than humans who answered the questions on the web. |
Copied to clipboard
| Challenge: | Cognitive science has long promoted the formation of mental models as central to understanding and question-answering. |
| Approach: | They train a new model, DREAM, to answer questions that elaborate the scenes that situated questions are about and then provide those elaborations as additional context to a question-answering (QA) model. |
| Outcome: | The proposed model is able to create better scene elaborations than a representative state-of-the-art, zero-shot model. |
Copied to clipboard
| Challenge: | Existing methods for QA are hampered by increased training costs . current methods suffer significant performance degradation when applied to out-of-domain examples. |
| Approach: | They propose a method that combines prompting methods and linear probing with fine-tuning strategy, which does not entail additional cost. |
| Outcome: | The proposed method outperforms state-of-the-art baselines with an average increase in F1 score of 4.5%-7.9%. |
Copied to clipboard
| Challenge: | Extractive question answering models are trained to predict start and end positions of answers . recent QA models outperform humans in some datasets due to their simplicity and effectiveness. |
| Approach: | They propose to use prior distribution of answer positions as a bias model to reduce position bias. |
| Outcome: | The proposed model outperforms BERT from 37.48% to 81.64% when trained on a biased SQUAD dataset. |
Copied to clipboard
| Challenge: | UnSeenTimeQA is a data contamination-free time-sensitive question-answering benchmark. |
| Approach: | They propose a data contamination-free time-sensitive question-answering benchmark that avoids web-searchable queries grounded in the real world. |
| Outcome: | The proposed benchmark avoids web-searchable queries grounded in the real world and enables on-demand generation of new samples, mitigating the risk of data leakage. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) performance on medical multiplechoice question (MCQ) benchmarks have stimulated interest from healthcare providers and patients globally. |
| Approach: | They introduce AfriMed-QA, the first largescale Pan-African English multi-specialty medical Question-Answering (QA) dataset, with 15,000 questions sourced from over 60 medical schools across 16 countries. |
| Outcome: | The proposed model outperforms other models in the medical field and is compared with other models. |
Copied to clipboard
| Challenge: | Despite the importance of datasets for natural language understanding, there has been little attention on crowdsourcing methods for collecting datasets. |
| Approach: | They compare the effectiveness of crowdsourcing methods for boosting NLU example difficulty with training crowdworkers instead of expert judgments. |
| Outcome: | The proposed method is ineffective for boosting NLU example difficulty, but it is not effective for training crowdworkers and qualifying workers based on expert judgments. |
Copied to clipboard
| Challenge: | Existing work on augmenting question answering models with external knowledge (e.g., knowledge graphs) lacks transparency into the model’s prediction rationale. |
| Approach: | They propose a knowledge-aware approach that equips pre-trained language models with a multi-hop relational reasoning module that performs multi-relational reasoning over subgraphs extracted from external knowledge graphs. |
| Outcome: | The proposed model performs multi-hop, multi-relational reasoning over subgraphs extracted from external knowledge graphs. |
Copied to clipboard
| Challenge: | Existing time-sensitive question answering models are limited for hard time-sensitive questions whose time qualifiers are implicit in the document. |
| Approach: | They propose a time-sensitive question answering framework that matches temporal events in documents with time qualifiers. |
| Outcome: | The proposed model outperforms baseline models for hard time-sensitive questions with 12.7% improvement in EM scores. |
Copied to clipboard
| Challenge: | Evidence-Based QA has proved insufficiently faithful with Large Language Models . a typical application of LLMs is in Evidence-based Question Answering (QA). |
| Approach: | They propose a data generation pipeline with automated data quality filters to fine-tune LLMs for better source quality and answer attributability. |
| Outcome: | The proposed model can synthesize high-quality training and testing data at scale. |
Copied to clipboard
| Challenge: | Reliable uncertainty quantification (UQ) is essential when employing large language models in high-risk domains such as clinical question answering (QA). |
| Approach: | They evaluate uncertainty estimation methods for clinical question answering using eleven clinical specialties and six question types. |
| Outcome: | The proposed method is based on behavioral features derived from reasoning-oriented models and examines conformal prediction as a complementary set-based approach. |
Copied to clipboard
| Challenge: | Existing data sets for document-level question answering are limited in their ability to detect short text and require multiple-sentence descriptive answers and opinions. |
| Approach: | They introduce a new data set with baseline methods for non-factoid long question answering . they compare BERT, RoBERTa, and Longformer models to establish baseline performances . |
| Outcome: | Experimental results show that Longformer outperforms the other architectures but human evaluations show that it is far behind the human upper bound. |
Copied to clipboard
| Challenge: | Recent advances in the field of language modeling have improved state-of-the-art results on many natural language processing tasks. |
| Approach: | They propose to use a French Question Answering Dataset to track progress of French Question answering models. |
| Outcome: | The proposed model achieves an F1 score of 92.2 and an exact match ratio of 82.1 on the test set. |
Copied to clipboard
| Challenge: | Developing a virtual assistant is crucial for supporting clients as it provides 24/7 assistance . factual questionanswering system is capable of handling all user queries . |
| Approach: | They propose a production-ready factual question answering system that combines local knowledge base search with generative, context-based QA. |
| Outcome: | The proposed system boosts local knowledge base retrieval by 23% . the system is language-agnostic and can be applied to any data domain . |
Copied to clipboard
| Challenge: | Existing pretrained language models have solved reading comprehension benchmarks, but datasets with information-seeking queries remain challenging. |
| Approach: | They analyze why answering information-seeking queries is more challenging . they manually annotate 800 unanswerable examples across six languages . |
| Outcome: | The proposed model outperforms human annotators on 800 unanswerable examples across six languages. |
Copied to clipboard
| Challenge: | a large number of indicators are difficult to measure, such as unemployment rate . a novel approach to measure socio-economic indicators with news events is proposed . |
| Approach: | They propose an event-centric indicator measure to extract news events from streaming news . they show strong correlations between ECIM values and representative indicators . |
| Outcome: | The proposed method is effective and correlated with several indicators . it is based on events reported in streaming news . |
Copied to clipboard
| Challenge: | Existing studies have focused on questions asked by experts, such as lawyers or legal scholars. |
| Approach: | They use a dataset to analyze laymen's legal questions paired with answers from lawyers and grounded to concrete law book paragraphs to find out what limitations exist. |
| Outcome: | The proposed system could help laymen in real situations without understanding law . the proposed system is based on 21k laymen’s legal questions paired with answers from lawyers and grounded to concrete law book paragraphs. |
Copied to clipboard
| Challenge: | Existing datasets focus on answerable questions or use automatically generated unanswerable questions that are easy to identify. |
| Approach: | They propose a dataset that combines the Stanford Question Answering Dataset with 50,000 unanswerable questions written by crowdworkers to look similar to answerable ones. |
| Outcome: | The proposed dataset looks similar to answerable questions on crowd-written questions . strong neural system that gets 86% F1 on SQuAD achieves only 66% F1. |
Copied to clipboard
| Challenge: | Multiple-choice question answering (MCQA) benchmarks show near-human accuracy . but a single accuracy score is a poor proxy for competence . |
| Approach: | They propose a medical multiple-choice question answering (MCQA) benchmark that augments three standard medical MCQA datasets with open-ended answers and systematically perturbed options. |
| Outcome: | The proposed benchmarks show that high MCQA accuracy masks low reliability . MCQ is the dominant paradigm for assessing medical knowledge in large language models . |
Copied to clipboard
| Challenge: | EconLogicQA requires models to discern and sequence multiple interconnected events, capturing the complexity of economic logics. |
| Approach: | They propose a benchmark to assess the sequential reasoning capabilities of large language models (LLMs) EconLogicQA requires models to discern and sequence multiple interconnected events, capturing the complexity of economic logics. |
| Outcome: | The proposed benchmark is based on a set of multi-event scenarios derived from economic articles and evaluates it across leading-edge LLMs. |
Copied to clipboard
| Challenge: | Neural ranking models require substantial amounts of relevance annotations, which is costly to scale. |
| Approach: | They propose to train a NR model with weak supervision instead of annotations . they use a structured overview of standard WS signals used for training a model . |
| Outcome: | The proposed approach reduces the cost of annotations by using weak supervision instead of a parametric model. |
Copied to clipboard
| Challenge: | Existing methods for solving geometric problems are limited due to lack of high-quality datasets and efficient neural solvers. |
| Approach: | They propose to annotate 2,518 geometric problems with richer types and greater difficulty using a benchmark dataset. |
| Outcome: | The proposed method improves the accuracy of automatic geometric problem solving to 66.09%. |
Copied to clipboard
| Challenge: | In-context examples can improve the performance of knowledge-rich tasks such as question answering by triggering a language model to surface information stored in its parametric knowledge. |
| Approach: | They propose to construct in-context example sets based on model's parametric knowledge by prompting models with 'unknown' examples. |
| Outcome: | The proposed model can perform better on in-context examples in three multi-answer question answering datasets, and prompting with ‘unknown’ examples decreases the performance. |
Copied to clipboard
| Challenge: | Existing methods to calibrate open setting machine reading systems fail to scale to these settings due to various scale limitations in practical settings. |
| Approach: | They propose to extend existing calibration approaches to calibrate open-domain question answering and claim verification systems to these settings. |
| Outcome: | The proposed calibration methods can selectively predict answers when question answering systems are posed with unanswerable or out-of-the-training distribution questions. |
Copied to clipboard
| Challenge: | a system that can show how its answers are implied by its own internal beliefs via a systematic chain of reasoning would allow better understanding of why a model produced the answer it did. |
| Approach: | They propose to combine a backward-chaining model with a verifier that checks that the model itself believes those premises through self-querying to generate multistep chains that are both faithful (the answer follows from the reasoning) |
| Outcome: | The proposed model generates chains that are faithful and truthful while maintaining answer accuracy. |
Copied to clipboard
| Challenge: | Traditionally, corpora are limited to arguments within the same sentence, and inter-sentential arguments are more challenging and have received less attention. |
| Approach: | They propose a question-answering approach to extract document-level event-argument structures by automating questions for each argument type an event may have. |
| Outcome: | The proposed model outperforms previous models and is especially beneficial to extract arguments that appear in different sentences than the event trigger. |
Copied to clipboard
| Challenge: | Existing work on QA explanation proposes to explain the answers with entailment trees composed of multiple enlargement steps. |
| Approach: | They propose a Module-based Entailment Tree GENeration framework that has multiple modules and a reasoning controller. |
| Outcome: | The proposed framework outperforms state-of-the-art models on the standard benchmark with only 9% of the parameters. |
Copied to clipboard
| Challenge: | Open-domain question answering (Open-QA) evaluations are criticized for the ambiguity in questions and the lack of semantic understanding in evaluators. |
| Approach: | They propose to examine the entailment relations of answers to identify more informative and more general system answers. |
| Outcome: | The proposed evaluations offer a much closer evaluation to human judgment on NaturalQuestions and TriviaQA while being learning-free. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used to answer factual, information-seeking questions (ISQs). |
| Approach: | They propose to use a dataset to evaluate large language models to generate human-like text on ISQs in two languages, English and Farsi, and then use it to evaluate nine LLMs. |
| Outcome: | The proposed dataset shows that accuracy drops by 25% when models encounter misleading yet factual hints. |
Copied to clipboard
| Challenge: | Recent work on Open Domain Question Answering has shown that there is a large discrepancy in model performance between novel test questions and those that largely overlap with training questions. |
| Approach: | They introduce and annotate questions according to three categories that measure training set overlap, compositional generalization, and novel-entity generalization. |
| Outcome: | The proposed models perform better on established datasets and lower on comp-gen/novel-entity questions than on the full test set. |
Copied to clipboard
| Challenge: | Existing approaches struggle with consistency across multiple languages and multi-size input scenarios. |
| Approach: | They propose a cross-lingual training framework that leverages multi-task learning to enhance cross-linguistic consistency and ranking stability. |
| Outcome: | The proposed training framework outperforms competitors on various input sizes and architectures. |
Copied to clipboard
| Challenge: | Existing DS-QA models ignore rich information contained in other paragraphs and are noisy . Existing systems rely on pre-identified relevant texts, which do not always exist in real-world QA scenarios. |
| Approach: | They propose a model which uses a paragraph selector to filter out noisy paragraphs and a reader to extract the correct answer from denoised paragraphs. |
| Outcome: | The proposed model can capture useful information from noisy data and achieve significant improvements on open domain question answering. |
Copied to clipboard
| Challenge: | Existing language agent systems struggle with costly data reliance and need multiple models for multiple functions. |
| Approach: | They propose an automatic agent learning framework for QA that synthesizes planning trajectories without human intervention. |
| Outcome: | The proposed framework outperforms existing models on question-answering tasks with a division-of-labor strategy. |
Copied to clipboard
| Challenge: | NLP models learn social biases, but little work has been done on how these biase manifest in outputs for applied tasks like question answering (QA). |
| Approach: | They propose a dataset that highlights attested social biases against people belonging to protected classes along nine social dimensions relevant for U.S. English-speaking contexts. |
| Outcome: | The proposed dataset highlights attested social biases against people belonging to protected classes along nine social dimensions relevant for U.S. English-speaking contexts. |
Copied to clipboard
| Challenge: | Recent years have brought about very fast developments in Natural Language Processing (NLP), but many other languages are overlooked due to limited resources. |
| Approach: | They propose to repurpose a multilingual BELEBELE dataset for a task of extractive QA in the style of machine reading comprehension. |
| Outcome: | The proposed approach could be used to extract QA in the style of machine reading comprehension. |
Copied to clipboard
| Challenge: | Question answering (QA) tasks have been posed using a variety of formats . a new study aims to develop specialized QA models that can be used to train QA systems . |
| Approach: | They build a pre-trained question answering model that performs well across 19 QA datasets . they argue that format-specialized models can limit the ability to teach reasoning . |
| Outcome: | a new model that trains on QA datasets performs on par with 8 models trained on individual datasets . a single model that trained on UNIFIEDQA performs well on 19 QA data . |
Copied to clipboard
| Challenge: | Existing work on calibration focuses on model confidence, such as the max probability of the predicted class. |
| Approach: | They propose a calibration method which estimates whether model correctly predicts answer for each question. |
| Outcome: | The proposed calibration method achieves 5-10% gains on reading comprehension benchmarks. |
Copied to clipboard
| Challenge: | Existing deep learning methods for answer selection are not feature engineering or expensive external resources. |
| Approach: | They propose to use deep learning methods to analyze and predict answer quality . they use a set of candidate answers to identify which of the candidates answers the question correctly. |
| Outcome: | The proposed methods produce impressive performance without feature engineering or expensive external resources. |
Copied to clipboard
| Challenge: | Medical board exams or general clinical questions do not capture the complexity of real clinical cases. |
| Approach: | They construct two datasets that are structured as multiple-choice question-answering tasks accompanied by expert-written explanations. |
| Outcome: | The proposed datasets are harder than previous benchmarks. |
Copied to clipboard
| Challenge: | Existing efficient test-time scaling methods introduce budget constraints or early stop mechanisms to avoid overthinking for straightforward questions but add human bias to the reasoning process. |
| Approach: | They propose a framework that dynamically adapts reasoning depth based on question complexity. |
| Outcome: | Experimental results show that the proposed framework achieves higher accuracy than baseline methods and reduces computational overhead by up to 25.2%. |
Copied to clipboard
| Challenge: | a lack of diverse and comprehensive question-answering datasets exists in under-resourced languages like Bangla. |
| Approach: | They propose a reading comprehension-based Bangla question-answering dataset . the dataset includes answerable and unanswerable questions covering four categories of questions . |
| Outcome: | The proposed dataset shows that it performs well as a training resource in high-resource languages. |
Copied to clipboard
| Challenge: | a product-related community question answering platform is widely employed in many E-commerce sites . however, the misinformation in the answers on those platforms poses unprecedented challenges for users to obtain reliable and truthful product information. |
| Approach: | They propose a large scale fact checking dataset from product question answering forums to predict the answer veracity . each answer is accompanied by its veraity label and associated evidence sentences . |
| Outcome: | The proposed model outperforms baselines on the question veracity prediction task. |
Copied to clipboard
| Challenge: | Several modern machine-learning based NLP systems can provide a confidence score with their output predictions. |
| Approach: | They propose a general calibration scheme for output entities of interest in NLP applications that can be used to calibrate confidence scores. |
| Outcome: | The proposed calibration scheme outperforms current calibration techniques for Named Entity Recognition, Part-of-speech tagging and Question Answering systems. |
Copied to clipboard
| Challenge: | A prominent challenge for language understanding systems is the ability to answer implicit reasoning questions where the evidence for answering the question is not mentioned explicitly. |
| Approach: | They propose to decouple inference of reasoning steps from execution by evaluating models of implicit relation inference. |
| Outcome: | The proposed model fails on the implicit reasoning QA task, but infers implicit relations . the proposed model is compared with other models that fail on the same task . |
Copied to clipboard
| Challenge: | Extractive QA models have shown promising performance in predicting the correct answer to a given question. |
| Approach: | They propose a BLANC-based context prediction task that learns the context prediction tasks. |
| Outcome: | The proposed model outperforms the state-of-the-art models on reading comprehension and hotpotQA. |
Copied to clipboard
| Challenge: | EASE is a diagnostic tool for Visual Question Answering (VQA) it quantifies the difficulty of an image, question sample. |
| Approach: | They propose a diagnostic tool which quantifies the difficulty of an image, question sample. |
| Outcome: | The proposed tool can be used to select the most-informative samples for training/fine-tuning. |
Copied to clipboard
| Challenge: | Prior work has investigated the ability of LLMs to abstain from answering context-dependent questions when provided insufficient or inconsistent context is provided. |
| Approach: | They propose to improve abstention when provided insufficient or incorrect context . they probed the ability of LLMs to abstain from answering context-dependent science questions . |
| Outcome: | The proposed models abstain from answering science questions when provided insufficient or incorrect context. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable performances in general domains and are now extending into the expert domain of law. |
| Approach: | They propose a Korean Benchmark for Legal EXplainable QA (KoBLEX) that evaluates provision-grounded, multi-hop legal reasoning. |
| Outcome: | The proposed method outperforms baselines and shows a high correlation with human judgments. |
Copied to clipboard
| Challenge: | Existing machine reading comprehension (MRC) models do not scale effectively to real-world applications like web-level information retrieval and question answering (QA). |
| Approach: | They propose a method that reframes existing machine reading comprehension (MRC) datasets as interactive, partially observable environments. |
| Outcome: | The proposed method "occludes" the majority of a document’s text and adds context-sensitive commands that reveal "glimpses" of the hidden text to a model. |
Copied to clipboard
| Challenge: | Large Language Models are a powerful tool for medical research, but the data is a bottleneck. |
| Approach: | They propose to use the largest ever medical Question Answering dataset with 26 Million QA pairs as a fine-tuning data for training large language models. |
| Outcome: | The proposed dataset demonstrates that it can be used to train large language models and improves zero-shot performance on other datasets. |
Copied to clipboard
| Challenge: | Existing studies on multi-hop question answering employ specific methods regardless of question types . complexity of multihop question answerrs often exceeds knowledge boundaries of LLMs . |
| Approach: | They propose a framework that uses chain-of-thought prompting to prompt LLMs to answer multi-hop questions. |
| Outcome: | The proposed framework outperforms baseline models in multi-hop QA scenarios. |
Copied to clipboard
| Challenge: | Existing studies on question answer matching focus on formal text . however, there exists many scenarios where the QA text is informal . |
| Approach: | They propose a novel QA matching approach using informal text from a product review site. |
| Outcome: | The proposed approach improves word-level and sentence-level attentions for solving the noisy problem in the informal text. |
Copied to clipboard
| Challenge: | Existing models generate erroneous information and evaluations fail to assess factual correctness of models. |
| Approach: | They propose to use MoleculeQA to evaluate molecular factual correctness in large language models by organizing molecules into a taxonomy and building QA pairs through human and LLM efforts. |
| Outcome: | The proposed model improves the factual correctness of generated information and enables the development of new models. |
Copied to clipboard
| Challenge: | Existing approaches to assess whether a given context contains sufficient information fail on factual questions. |
| Approach: | They propose a framework that asks a model to reason about what information is missing . this framework generates more accurate sufficiency judgments while articulating any information gaps . |
| Outcome: | The proposed framework produces more accurate sufficiency judgments while clearly articulating any information gaps. |
Copied to clipboard
| Challenge: | Existing approaches to answer questions using large language models lack the ability to faithfully follow the intermediate reasoning steps from the known premises to the answer. |
| Approach: | They propose a faithful question-answering task that uses a Monte-Carlo planning algorithm to produce faithful reasoning steps from the known premises to the answer. |
| Outcome: | The proposed task can produce valid and faithful reasoning steps compared with large language models with a much smaller model size. |
Copied to clipboard
| Challenge: | Existing work on decompositions of complex questions has focused on multi-step reasoning . but, in machine reading, it is unclear when decomposing is helpful . |
| Approach: | They conduct experiments on decompositions in machine reading to unify recent work . they find that decomposing complex questions can be helpful in zero or limited-data settings . |
| Outcome: | The proposed model can learn decompositions implicitly even with limited data, the study shows . the results are consistent with previous work on decomposing complex questions . |
Copied to clipboard
| Challenge: | Using neural question answering models, our system generates answer candidates and then combines loopy belief propagation with local search to find full puzzle solutions. |
| Approach: | They propose a new approach to automatically solving crossword puzzles that uses neural question answering models and loopy belief propagation with local search to find full puzzle solutions. |
| Outcome: | The proposed system outperforms even the best human solvers and can solve crosswords from a wide range of domains with perfect accuracy. |
Copied to clipboard
| Challenge: | False. a free-form question answering dataset can serve as a useful research benchmark for source code comprehension. |
| Approach: | They propose a free-form question answering dataset for source code comprehension . they implement syntactic rules and semantic analysis to transform code comments into question-answer pairs. |
| Outcome: | The proposed dataset can serve as a useful research benchmark for source code comprehension. |
Copied to clipboard
| Challenge: | Existing studies on adapting large language models to perform a variety of tasks in high-stakes domains such as healthcare lack understanding of the extent and contributing factors that allow them to recall relevant knowledge and combine it with presented information. |
| Approach: | They propose to use multiple choice and abstractive question answering to investigate the extent and contributing factors that allow LLMs to recall relevant knowledge and combine it with presented information in the clinical and biomedical domain. |
| Outcome: | The proposed models perform better on 22 datasets in three generalist and three specialist biomedical sub-domains, and show that they can generalise to unseen sub- domains. |
Copied to clipboard
| Challenge: | a system that finds the strongest supporting evidence for a given answer is proposed . a study using passage-based question-answering (QA) shows that agents select evidence that generalizes . |
| Approach: | They propose a system that finds the strongest supporting evidence for a given answer . they use passage-based question-answering (QA) as a testbed to train evidence agents . |
| Outcome: | The proposed system improves QA in a robust manner by using agent-selected evidence. |
Copied to clipboard
| Challenge: | Current retrieval-augmented generation systems struggle when retrieval models fail to rank the most relevant documents . existing extractive methods reduce latency but rely on independent, non-adaptive sentence selection . |
| Approach: | They introduce an extractive context compression framework that enhances retrieval-augmented generation in question answering. |
| Outcome: | EXIT surpasses existing compression methods and uncompressed baselines in QA accuracy . the framework reduces inference time and token count while preserving contextual dependencies . |
Copied to clipboard
| Challenge: | Existing datasets for reading comprehension have deterministic answers, but questions in the real world do not always have definite answers. |
| Approach: | They propose a Question Answering (QA) dataset that contains complex questions with conditional answers. |
| Outcome: | The proposed dataset will motivate further research in answering complex questions over long documents. |
Copied to clipboard
| Challenge: | Existing approaches to answer multiple-choice questions with no supporting documents are poor performance. |
| Approach: | They propose a method which can be used to semantically rank documents extracted from Wikipedia . they propose 'semantic ranking' method that latently learns to rank documents by their importance . |
| Outcome: | The proposed model achieves state-of-the-art accuracy on two datasets: ARC Easy and Challenge. |
Copied to clipboard
| Challenge: | Existing literature observes bias in question answering (QA) models, but there is no method to mitigate it. |
| Approach: | They propose an approach to mitigate the bias of question answering models by observing the influence of a query instance on another instance. |
| Outcome: | The proposed method reduces bias level in all 9 bias categories while maintaining comparable QA accuracy. |
Copied to clipboard
| Challenge: | Existing annotations for other NLP tasks are used to generate domain-specific large-scale question answering (QA) datasets. |
| Approach: | They propose to re-purpose existing annotations for other NLP tasks by generating a large-scale question answering corpus using 1 million questions-logical form and 400,000+ question-answer evidence pairs. |
| Outcome: | The proposed model can be trained to learn domain-specific large-scale question answering (QA) datasets. |
Copied to clipboard
| Challenge: | PubMedQA is a biomedical question answering dataset based on PubMed abstracts . 68.1% accuracy is achieved, compared to single human performance of 78.0% . |
| Approach: | They propose a biomedical question answering dataset from PubMed abstracts . the dataset is annotated by experts and has 1k instances of QA . |
| Outcome: | The proposed model achieves 68.1% accuracy compared to human performance of 78.0% and majority-baseline of 55.2%. |
Copied to clipboard
| Challenge: | Current long-context large language models lack citations to support their responses, making verification difficult due to potential hallucinations. |
| Approach: | They propose to use off-the-shelf LLMs to automatically construct long-context QA instances with precise sentence-level citations and leverage this pipeline to construct a large-scale SFT dataset for LQAC. |
| Outcome: | The proposed pipeline can generate responses with fine-grained citations on the fly, surpassing existing models including GPT-4o. |
Copied to clipboard
| Challenge: | Open-domain and multi-hop QA is an important problem for both humans and computers. |
| Approach: | They propose a gamified interface where a human answers complex questions with access to traditional and modern search tools. |
| Outcome: | The proposed interface compares human queries to state-of-the-art QA models . human queries can improve the accuracy of existing systems, the authors argue . |
Copied to clipboard
| Challenge: | Visual Question Answering (VQA) models are prone to learn the shortcut solution formed by dataset biases rather than the intended solution. |
| Approach: | They propose a dataset that considers varying types of shortcuts by constructing different distribution shifts in multiple OOD test sets. |
| Outcome: | The proposed dataset considers varying types of shortcuts by constructing different distribution shifts in multiple OOD test sets. |
Copied to clipboard
| Challenge: | Existing methods for processing large textual content face insufficient adaptation to task-specific needs and missing multi-segmentation relationships. |
| Approach: | They propose a question then reflection memory mechanism which integrates a dual-structured memory pool and a structured graph guidance to facilitate a reflective trial-and-error approach for navigating and identifying relevant segments. |
| Outcome: | The proposed model achieves superior performance on multiple-choice questions and multi-doc QA. |
Copied to clipboard
| Challenge: | Existing approaches to answer selection are limited in domains with limited labeled data. |
| Approach: | They propose a Knowledge-aware Attentive Network framework for cross-domain answer selection that uses the knowledge base as a bridge to enable knowledge transfer from the source domain to the target domain. |
| Outcome: | The proposed model outperforms strong competitors by a noticeable margin in cross-domain answer selection. |
Copied to clipboard
| Challenge: | Ambiguous questions are a challenge for Question Answering models as they require answers that cover multiple interpretations of the original query. |
| Approach: | They aim to investigate whether model/data scaling improves the answers’ quality and whether automated metrics align with human judgment. |
| Outcome: | The proposed models can generate long-form answers that combine conflicting information and provide valuable insights into the limitations of the current approaches. |
Copied to clipboard
| Challenge: | Existing machine reading comprehension tasks lack interactive information-seeking component of comprehension. |
| Approach: | They propose a question-asking task that asks questions in a text-based environment . they propose QAit, which uses a game generator to build models that include deep reinforcement learning agents. |
| Outcome: | The proposed task poses questions about existence, location, and attributes of objects found in environment. |
Copied to clipboard
| Challenge: | Multi-hop textual question answering requires combining information from multiple sentences. |
| Approach: | They propose a model that explicitly identifies the knowledge gap between a key span in the provided knowledge and the answer choices. |
| Outcome: | The proposed model outperforms existing models on the OpenBookQA dataset. |
Copied to clipboard
| Challenge: | Existing approaches do not fully exploit the interdependency between document and query. |
| Approach: | They propose a novel dependent gated reading bidirectional GRU network to efficiently model the relationship between the document and the query during encoding and decision making. |
| Outcome: | The proposed model performs well on machine comprehension benchmarks such as the Children’s Book Test and Who DiD What. |
Copied to clipboard
| Challenge: | Existing models fail to answer a large portion of sub-questions . Existing systems have achieved super-human performance . |
| Approach: | They propose to use a neural decomposition model to generate sub-questions for a multi-hop question and extract the corresponding sub-answers. |
| Outcome: | The proposed model is based on a hotpotQA dataset with a multi-hop question and sub-answers. |
Copied to clipboard
| Challenge: | Temporal question answering (QA) is a complex task that requires reasoning over facts asserting time intervals of events. |
| Approach: | They propose a temporal fact extraction technique that helps QA when it fails to retrieve temporal facts from the KB. |
| Outcome: | The proposed technique can extract temporal facts that failed to get retrieved from the KB without additional training cost. |
Copied to clipboard
| Challenge: | Existing causal question answering datasets are relatively small and only include one type of causal question. |
| Approach: | They construct a benchmark corpus of 1.1 million causal questions with answers . they use a typology derived from a data-driven, manual analysis of QA datasets . |
| Outcome: | The proposed model achieves a ROUGE-L F1 score of 0.48 on the new QA benchmark. |
Copied to clipboard
| Challenge: | End-to-end neural networks excel at answering natural language questions but fail on complex ones . a proposed framework for question parsing and execution on textual QA is designed to combine the strengths of neural and symbolic methods. |
| Approach: | They propose a framework for question parsing and execution on textual QA . they parse questions into an intermediate representation and use deterministic rules to translate them . |
| Outcome: | The proposed framework outperforms existing methods in supervised, few-shot, and zero-shot settings while preserving its underlying reasoning process. |
Copied to clipboard
| Challenge: | Existing knowledge based question answering systems are trained based on labeled reasoning paths, which hinder their performance. |
| Approach: | They propose a KBQA system which leverages multiple reasoning paths’ information and only requires labeled answer as supervision. |
| Outcome: | The proposed system can leverage multiple reasoning paths’ information and only requires labeled answer as supervision. |
Copied to clipboard
| Challenge: | Existing question answering systems rely on pre-selected and annotated evidence documents, thus making them inadequate for addressing novel questions. |
| Approach: | They propose to use the common retrieve-then-read QA pipeline and PubMed as a trustworthy collection of medical research documents to answer health questions from three diverse datasets. |
| Outcome: | The proposed approach improves the macro F1 score by 10% by utilizing the common retrieve-then-read QA pipeline and PubMed as a trustworthy collection of medical research documents. |
Copied to clipboard
| Challenge: | Existing methods for knowledge base question answering ignore subtle inter-relationships between the question and the KB. |
| Approach: | They propose to model the two-way flow of interactions between questions and KBs using a bidirectional attentive memory network. |
| Outcome: | The proposed method outperforms existing methods on the WebQuestions benchmark and offers better interpretability compared to baselines. |
Copied to clipboard
| Challenge: | Open-domain complex question-answering systems face challenges in retrieving and reasoning over information that addresses multifaceted queries. |
| Approach: | They propose a method that leverages large language models to guide a Neighborhood Aware Retrieval process. |
| Outcome: | The proposed approach outperforms retrieve-and-reason baselines on two complex QA datasets. |
Copied to clipboard
| Challenge: | Existing Question-Answering (QA) datasets contain unanswerable questions . however, their treatment in QA systems remains primitive . |
| Approach: | They propose a framework that provides answers based on presupposition failure over oracle behavior of existing QA systems. |
| Outcome: | The proposed system provides responses based on presupposition failure over oracle behavior of existing QA systems. |
Copied to clipboard
| Challenge: | Pre-trained language models (PLMs) are often easily deceived by tricky questions such as “How many eyes does the sun have?” . |
| Approach: | They annotate a FalseQA dataset containing 2365 human-written FPQs and find that PLMs are capable of discriminating FPqs by fine-tuning on moderate numbers. |
| Outcome: | The proposed model can discriminate on FPQs by fine-tuning on moderate numbers of examples and generate reasonable explanations for false premise questions. |
Copied to clipboard
| Challenge: | Language embeddings have been shown to have stereotyping biases, but how these biase affecting downstream question answering models remains unexplored. |
| Approach: | They propose a general framework to probe biases through underspecified questions by building minimal context and building minimal questions. |
| Outcome: | The proposed framework isolates two types of reasoning errors and identifies stereotyping biases in gender, nationality, ethnicity, and religion classes. |
Copied to clipboard
| Challenge: | Recent work has shown that good performance on a dataset might not correlate well with human’s expectations from models that “understand” language. |
| Approach: | They propose to train a top performing multiple choice question answering model against expectations from models that "understand" language. |
| Outcome: | The proposed training paradigm leads to a model that performs on par with the original model while better satisfying our expectations. |
Copied to clipboard
| Challenge: | Recent adaptive retrieval methods integrate LLMs’ intrinsic knowledge with external information appealing to LLM self-knowledge, but they often neglect efficiency evaluations and comparisons with uncertainty estimation techniques. |
| Approach: | They propose to integrate LLMs’ intrinsic knowledge with external information appealing to LLM self-knowledge but neglect efficiency evaluations and comparisons with uncertainty estimation techniques. |
| Outcome: | The proposed methods outperform complex pipelines in terms of efficiency and self-knowledge while maintaining comparable QA performance. |
Copied to clipboard
| Challenge: | Recent question answering systems perform well on benchmark datasets, but are not always well-calibrated to spot spurious answers under distribution shifts. |
| Approach: | They propose to use natural language inference to verify whether answers are correct . they leverage large pre-trained models and recent prior datasets to construct powerful question conversion and decontextualization modules. |
| Outcome: | The proposed approach improves the confidence estimation of a QA model across different domains, evaluated in a selective QA setting. |
Copied to clipboard
| Challenge: | Existing evaluation methods rely on rule-based matching with shallow semantic understanding or adopt LLM-as-a-Judge approaches that incur high cost and latency while offering limited error interpretability. |
| Approach: | They propose a curriculum learning based hierarchical framework for QA task evaluation that supports quick scoring and fine-grained error analysis. |
| Outcome: | The proposed framework outperforms baseline methods on quick scoring and error analysis tasks while being 25 faster. |
Copied to clipboard
| Challenge: | AdvisorQA aims to improve LLMs’ capability to offer advice for deeply subjective concerns, utilizing the LifeProTips Reddit forum. |
| Approach: | They propose a dataset to train LLMs' ability to offer advice for deeply subjective concerns, utilizing the LifeProTips Reddit forum. |
| Outcome: | The proposed model improves usefulness through automatic metric, GPT-4 and human evaluations, and expands independent evaluation axis to include harmlessness. |
Copied to clipboard
| Challenge: | Existing QA systems for question answering are limited by the availability of annotated datasets. |
| Approach: | They propose a dataset for question-answering that extracts information from multiple parts of text . they propose QA-based multi-span neural architecture that captures relevance among multiple answer spans . |
| Outcome: | The proposed model outperforms state-of-the-art QA models in this multi-span QA setting. |
Copied to clipboard
| Challenge: | Existing question answering datasets provide extractive or short answers, but less attention has been paid to open-ended questions that require explanations. |
| Approach: | They present a large-scale corpus for long form question answering . they use a Reddit forum to provide elaborate answers to open-ended questions . |
| Outcome: | The proposed model outperforms Seq2Seq, language modeling, and other models in human evaluations. |
Copied to clipboard
| Challenge: | Recent research has shown that self-citing large language models (LLMs) fail to faithfully reflect their context usage throughout the generation process. |
| Approach: | They propose a plug-and-play approach using model internals for faithful answer attribution in RAG applications that detects context-sensitive answer tokens and pairs them with retrieved documents contributing to their prediction. |
| Outcome: | The proposed approach achieves citation quality and efficiency comparable to self-citation while allowing for a finer-grained control of attribution parameters. |
Copied to clipboard
| Challenge: | Existing studies on controversy define it based on vague assumptions of its relation to sentiment . experimental results show controversy detection is essential and challenging . |
| Approach: | They propose a question-answering dataset that defines content controversy by user perception . they show controversy detection is essential and challenging . |
| Outcome: | The proposed dataset defines controversy by user perception, i.e., votes from plenty of users. |
Copied to clipboard
| Challenge: | Empirical results show that AMATA outperforms baseline approaches, knowledge-augmented frameworks, and LLMs on knowledge-intensive QA benchmarks. |
| Approach: | They propose an Adaptive Multi-Agent Trajectory Alignment framework that integrates external knowledge to improve response interpretability and factual grounding. |
| Outcome: | The proposed framework outperforms baseline approaches, knowledge-augmented frameworks, and LLM-based trajectory systems on five established knowledge-intensive QA benchmarks. |
Copied to clipboard
| Challenge: | Modern systems for multi-hop question answering (QA) break questions into a sequence of reasoning steps, termed chain-of-thought (CoT) Often, multiple chains are sampled and aggregated, but the intermediate steps themselves are discarded. |
| Approach: | They propose a method which prompts large language models to meta-reason over multiple chains of thought rather than aggregate their answers. |
| Outcome: | The proposed approach outperforms baselines on 7 multi-hop QA datasets. |
Copied to clipboard
| Challenge: | Existing studies have focused on the spatial reasoning capabilities of modern language models (LMs) however, there has been limited research into the spatial thinking capabilities of LMs. |
| Approach: | They propose a question-answering (QA) benchmark for spatial reasoning on natural language text which contains more realistic spatial phenomena not covered by prior work. |
| Outcome: | The proposed method significantly improves LMs' ability on spatial understanding, which in turn helps solve two external datasets, bAbI, and boolQ. |
Copied to clipboard
| Challenge: | Existing information-seeking question answering datasets do not perform well on answering these questions . existing models that do well on other QA tasks do not do well answering these tasks . |
| Approach: | They present a dataset of 5049 questions over 1585 NLP papers . they use a question-seeking QA model that seeks information in the full text . |
| Outcome: | The proposed dataset underperforms existing models on other QA tasks by 27 F1 points . the focus is on document-grounded, information-seeking QA . |
Copied to clipboard
| Challenge: | Recent work has attempted to improve extractive QA performance by enriching the dataset with unanswerable questions. |
| Approach: | They build an out-of-domain corpus of competitive and non-competitive questions . they compare the results with the results of the Recognizing Textual Entailments task . |
| Outcome: | The proposed model fails even in the case of simpler questions . the proposed model can be used to address more realistic situations in reading comprehension . |
Copied to clipboard
| Challenge: | Experimental results show that our approach can effectively improve the performance of both the policy model and the reward model. |
| Approach: | They propose to use Monte Carlo Tree Search for both policy model improvement and reward model improvement to bridge it to more subtle open-domain question answering. |
| Outcome: | The proposed approach surpasses existing methods for annotation and training data with fewer data points and achieves better performance in test-time scaling strategies. |
Copied to clipboard
| Challenge: | S-MedQA is an English question-answering dataset designed for benchmarking large language models in fine-grained clinical specialties. |
| Approach: | They propose to use an English medical question-answering dataset to benchmark large language models in clinical specialties. |
| Outcome: | The proposed dataset is designed to benchmark large language models in medical specialties. |
Copied to clipboard
| Challenge: | Existing models for natural language understanding are limited to processing only a few hundred words at a time. |
| Approach: | They propose a dataset with context passages in English that have an average length of 5,000 tokens. |
| Outcome: | a new dataset with long-text comprehension questions is used to test models on long-document comprehension . the questions are validated by contributors who have read the entire passage, not just excerpts . only half of the questions can be answered by annotators working under tight time constraints . |
Copied to clipboard
| Challenge: | Extractive question answering models are reliant on annotations of answer-spans in the corresponding passages. |
| Approach: | They propose a method that auto-encodes a question and generates corresponding questions from it. |
| Outcome: | The proposed method performs well in a zero-shot setting and can provide an additional loss to boost performance for extractive question answering (EQA). |
Copied to clipboard
| Challenge: | Existing approaches to model long-range dependencies in text are limited to 512 tokens . however, the amount of compute in attention depends quadratically on the number of tokens in an input text passage. |
| Approach: | They propose a technique that summarises text into a memory table to be used in a second read of the text. |
| Outcome: | The proposed method outperforms models of comparable size on several question answering datasets and sets a new state of the art on the NarrativeQA task, with questions about entire books. |
Copied to clipboard
| Challenge: | Existing knowledge editing methods overlook interplay with pre-existing knowledge, leading to inconsistent edit propagation. |
| Approach: | stepKE integrates edited and existing knowledge for coherent multi-hop reasoning . stepKE decomposes multi-step questions into sequential single-hop sub-questions . |
| Outcome: | Experiments show that StepKE generates more accurate and consistent responses than baselines. |
Copied to clipboard
| Challenge: | Non-factoid (NF) question answering is challenging to evaluate due to diverse potential answers and no objective criterion. |
| Approach: | They propose a listwise NFQA evaluation approach that uses Large Language Models to rank candidate answers in a descending list of reference answers sorted by descending quality. |
| Outcome: | The proposed method has higher correlations with human annotations than standard methods. |
Copied to clipboard
| Challenge: | Attributed Question Answering models are not yet leveraged to enhance their essential capabilities, including evidence identification, cross-source relation recognition and anti-distraction reasoning. |
| Approach: | They propose a progressive progressive curriculum learning approach that optimizes both encoder-decoder and decoder-only AQA models. |
| Outcome: | The proposed approach improves both encoder-decoder and decoder-only AQA models on the quotesum benchmark. |
Copied to clipboard
| Challenge: | Existing metrics for evaluating the quality of automatically generated questions are expensive and penalise valid questions that may not have high lexical or semantic similarity to the reference questions. |
| Approach: | They propose a question-answering and span scorer metric based on the answerability of the candidate question given the context. |
| Outcome: | The proposed metric has higher correlation with human judgment without relying on the reference question. |
Copied to clipboard
| Challenge: | Community Question Answering websites are becoming popular and useful source of information for users. |
| Approach: | They propose to use community question answering forum to detect similar questions . they use question-answering similarity task to provide correct answers . |
| Outcome: | The proposed framework provides the first framework on the evaluation of similar questions and question-answering detection on a multi-domain corpora. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) often struggle with question answering due to hallucinated answers. |
| Approach: | They propose a multilingual QA dataset with evergreen labels that can be used to evaluate and train large language models. |
| Outcome: | The proposed model performs well on 12 modern LLMs and EG-E5 classifiers. |
Copied to clipboard
| Challenge: | Existing work on multi-domain, multi-lingual question answering is limited to the same language. |
| Approach: | They curate 500 articles in six different domains from the web and create question-answer pairs . they develop a deep learning based model for classifying an input question into coarse and finer categories . |
| Outcome: | The proposed model accuracies 90.12% and 80.30% for coarse and finer classes . the proposed model is the first attempt to create multi-domain, multi-lingual question answering evaluation involving English and Hindi. |
Copied to clipboard
| Challenge: | Existing question answering datasets lack diversity in gender, profession, and nationality. |
| Approach: | They focus on how well QA models generalize across demographic subsets . english-language QA datasets mostly ask about US men from a few professions - this is problematic because most English speakers are not from the US or UK . |
| Outcome: | The proposed model accuracy is lower for people based on gender, profession, and nationality, but there is more variation on professions (question topic) and question ambiguity. |
Copied to clipboard
| Challenge: | Recent research in interpretability of neural models has yielded numerous token attribution techniques, but it is hard to evaluate whether these explanations are faithful. |
| Approach: | They propose to use pairwise attributions to connect outputs to high-level model behavior to examine how well different attribution techniques align with this assumption on realistic counterfactuals in the case of reading comprehension (RC). |
| Outcome: | The proposed methods are better suited to RC than token-level attributions across different RC settings, and the best performance comes from a modification that was proposed to an existing pairwise attribution method. |
Copied to clipboard
| Challenge: | Existing datasets that focus on temporal knowledge are limited in size and lack comprehensive coverage of temporal information. |
| Approach: | They introduce a large-scale temporal question-answer-matching dataset . the new taxonomy categorizes questions as attributes, comparisons, and counting questions . |
| Outcome: | The proposed dataset surpasses existing benchmarks in scale and scope. |
Copied to clipboard
| Challenge: | Reinforcement learning (RL) for large language models typically requires clear reward signals, which are often unavailable for open-ended (OE) questions where answer evaluation is ambiguous without scalable expert labeling. |
| Approach: | They propose a mixed-data approach to training large language models with varying reward clarity . they combine Multiple-choice questions (MCQs) with OE questions for which they use simpler, potentially noisy rewards such as Jaccard similarity or LLM-based evaluators. |
| Outcome: | The mixed-data approach improves medical question-answering performance across model scales. |
Copied to clipboard
| Challenge: | Existing open-domain question answering systems assume questions have a single welldefined answer. |
| Approach: | They propose an open-domain question answering task which involves finding every plausible answer and rewriting the question for each one to resolve the ambiguity. |
| Outcome: | The proposed task is based on a dataset covering 14,042 open-domain questions . it shows that strong models benefit from weakly supervised learning . |
Copied to clipboard
| Challenge: | Open Domain Multi-Hop Question Answering (ODMHQA) is one of the most challenging tasks in Natural Language Processing (NLP) |
| Approach: | They propose a mechanism that leverages the intrinsic capabilities of Large Language Models to judge whether the generated answers are off-topic. |
| Outcome: | The proposed method reduces the occurrence of off-topic answers by nearly 13%, improving the performance in Exact Match (EM) by nearly 3% compared to the baseline method without the Dr3 mechanism. |
Copied to clipboard
| Challenge: | Developing such datasets is important for the development and evaluation of Icelandic QA systems. |
| Approach: | They present the first extractive question answering dataset for Icelandic, Natural Questions in Icelandic. |
| Outcome: | The proposed dataset is a valuable resource for Icelandic which is being evaluated by a team of researchers. |
Copied to clipboard
| Challenge: | Identifying bridge phrases remains one of the challenges for multi-hop question answering . |
| Approach: | They propose an unsupervised method for the identification of bridge phrases in multi-hop question answering . they construct a graph of noun phrases from the question and available context . |
| Outcome: | The proposed method improves all downstream components in a multi-hop QA system. |
Copied to clipboard
| Challenge: | Taking the exam closed book, but having read the textbook, yields at best minor improvement (56%), suggesting that the PTLM may not have “understood” the textbook (or perhaps misundersttoo the questions). |
| Approach: | They propose to use pre-trained language models to answer questions from introductory college textbooks and hundreds of true/false statements based on review questions written by the authors. |
| Outcome: | The proposed task includes two college-level introductory texts in the social sciences (American Government 2e) and humanities (U.S. History). |
Copied to clipboard
| Challenge: | Existing approaches to QA over textual data are based on a "retrieve-then-generate" pipeline. |
| Approach: | They propose a "triple-level" labeling strategy that infers fine-grained labels and trains a re-ranker to improve relevance of retrieved triples. |
| Outcome: | The proposed pipeline improves on prior KGQA systems by 5.56% Exact Match. |
Copied to clipboard
| Challenge: | Social media is becoming an important realtime information source, especially during natural disasters and emergencies. |
| Approach: | They present a large-scale dataset for question answering over social media data . they gather tweets used by journalists and ask human annotators to write questions upon them . |
| Outcome: | The proposed dataset shows that neural models that perform well on formal texts are limited in their performance . the proposed model is still lagging behind human performance with a large margin . |
Copied to clipboard
| Challenge: | Question Answering models typically use retrieval and reasoning components to identify relevant information for reasoning. |
| Approach: | They propose a retrieval parameterization method that marginalizes unanswerable queries . they show that marginalization allows a model to mitigate false negatives in annotations . |
| Outcome: | The proposed model improves on two multi-document question answering datasets and shows that marginalization improves performance. |
Copied to clipboard
| Challenge: | Recent studies have shown that LLM-based EHR question answering is costly to deploy and does not leverage hierarchical structure of clinical data. |
| Approach: | They propose a Lorentzian model that embeds codes, visits, and questions in hyperbolic space and answers queries via geometry-consistent cross-attention with type-specific pointer heads. |
| Outcome: | The proposed model embeds codes, visits, and questions in hyperbolic space and answers queries via geometry-consistent cross-attention with type-specific pointer heads. |
Copied to clipboard
| Challenge: | Existing work suggests the appeals of incorporating explicit semantic representations into NLP . semi-structured natural language structures provide an intermediate meaning-capturing representation . |
| Approach: | They propose a semi-structured natural-language representation of textual information . they examine input and output linearization strategies and multitask learning . |
| Outcome: | The proposed model is based on pre-trained sequence-to-sequence language models . it is easy to use and can be used for downstream tasks that benefit from it . |
Copied to clipboard
| Challenge: | Existing studies show that children who excel at mindreading are more likely to be identified as popular by classmates and have reciprocated friendships. |
| Approach: | They propose to automate the scoring of mindreading ability in middle childhood and early adolescence using a new corpus of 11,311 question-answer pairs in English from 1,066 children aged from 7 to 14 . |
| Outcome: | The proposed scoring system is based on 11,311 question-answer pairs in English from 1,066 children aged from 7 to 14 . the results demonstrate the applicability of state-of-the-art NLP solutions to a new domain and task. |
Copied to clipboard
| Challenge: | Existing evaluation methods overlook the distinction between factoid and non-factoidic questions. |
| Approach: | They propose a method that distinguishes open-ended questions and ranks candidate answers . they propose QA requires longer answer statements and nuanced reasoning processes . |
| Outcome: | The proposed method better aligns with human annotations and offers more interpretable results. |
Copied to clipboard
| Challenge: | We introduce Fact-QA, an LLM-based evaluation metric to evaluate the factuality of predicted answers. |
| Approach: | They propose to use an open-source dataset to analyze logistics-related question-answer pairs in a logistics-based course. |
| Outcome: | The proposed approach performs close to humans on traditional metrics of textual similarity, but there is a significant gap between them and humans in terms of fact precision. |
Copied to clipboard
| Challenge: | Question answering models have access to two sources of knowledge during inference time: parametric knowledge and contextual knowledge. |
| Approach: | They propose a new paradigm in which QA models are trained to disentangle the two sources of knowledge. |
| Outcome: | The proposed model generates two answers for a given question based on parametric and contextual knowledge. |
Copied to clipboard
| Challenge: | Ambiguous questions have different answers depending on their interpretation and can take diverse forms. |
| Approach: | They propose a manually annotated temporally ambiguous QA dataset that captures temporal ambiguity and propose different search strategies based on disambiguate versions of the questions. |
| Outcome: | The proposed approach captures temporal ambiguity and provides non-search, competitive baselines for detecting temporal and few-shot ambiguities. |
Copied to clipboard
| Challenge: | Recent progress on factoid question answering (QA) does not easily transfer to the task of long-form QA where the goal is to generate detailed explanations. |
| Approach: | They propose a task that focuses on ambiguous factoid questions which have different correct answers depending on interpretation. |
| Outcome: | The proposed metric is reliable and demonstrates agreement between this metric and human judgments, and reveals a considerable gap between human performance and strong baselines. |
Copied to clipboard
| Challenge: | Existing QA systems do not strictly enforce cross-document synthesis or exploit the explicit inter-paper structure that links sources. |
| Approach: | They propose a pipeline methodology for constructing a multi-document academic QA dataset . they detect communities based on citation networks and leverage Large Language Models . |
| Outcome: | The proposed method generates QA pairs related to multi-document content automatically and forms coherent communities based on citation networks and large language models. |
Copied to clipboard
| Challenge: | Existing multi-hop question answering datasets do not provide a complete explanation for the reasoning process from the question to the answer. |
| Approach: | They propose a multi-hop question answering dataset that uses structured and unstructured data to test reasoning skills. |
| Outcome: | The proposed dataset ensures multi-hop reasoning while being challenging for multi-models. |
Copied to clipboard
| Challenge: | Existing contrastive methods that ignore the context of a large language model (LLM) fail to handle instances that vary in their amount of conflict, with static methods over-adjusting when conflict is absent. |
| Approach: | They propose a fine-grained, instance-level approach called AdaCAD which dynamically adjusts the degree of conflict based on the degree. |
| Outcome: | The proposed approach outperforms baselines and improves factuality of summaries by 6.19. |
Copied to clipboard
| Challenge: | Existing models struggle with producing answers that are frequently updated or from uncommon locations. |
| Approach: | They propose an open-retrieval QA dataset where systems must produce the correct answer given the context. |
| Outcome: | The proposed dataset shows that existing models struggle with producing answers that are frequently updated or from uncommon locations. |
Copied to clipboard
| Challenge: | interacting with a model for Visual Question Answering (VQA) quickly reveals that these models lack consistency. |
| Approach: | They propose a dataset, ConVQA, and metrics that enable quantitative evaluation of consistency in VQA. |
| Outcome: | The proposed data augmentation module improves the consistency of VQA models on the Con-VQA dataset and is a strong baseline for future research. |
Copied to clipboard
| Challenge: | SQA is an emerging application of NLP in the medical, geography, and legal domains. |
| Approach: | They propose a dataset of 1,981 scenarios and 4,110 multiple-choice questions in geography domain at high school level. |
| Outcome: | The proposed dataset consists of 1,981 scenarios and 4,110 multiple-choice questions in the geography domain at high school level. |
Copied to clipboard
| Challenge: | Existing approaches on semantic parsing suffer from exponential growth of logical form candidates and can hardly generalize to unseen data. |
| Approach: | They propose a unified semantic parser for question answering on KB and DB . they define the primitive as the essential element in their framework . |
| Outcome: | The proposed framework can predict logical forms by altering and composing top-ranked primitives with different operations. |
Copied to clipboard
| Challenge: | Recent research shows that relevant knowledge can provide useful context for commonsense tasks. |
| Approach: | They propose a method that learns to generate contextually relevant knowledge in response to given questions. |
| Outcome: | The proposed method shows consistent gains over 9 commonsense benchmarks. |
Copied to clipboard
| Challenge: | Recent research has shown that smaller language models can acquire substantial reasoning abilities when fine-tuned with reasoning exemplars crafted by a significantly larger teacher model. |
| Approach: | They propose to fine-tune several smaller model to generate programs that encode the required financial reasoning and calculations. |
| Outcome: | The proposed model outperforms the teacher model in the financial domain by adjusting the entity extraction for the specific data format. |
Copied to clipboard
| Challenge: | Existing methods for generating synthetic question answering corpora are not suitable for QA, but can be constructed from widely available natural text. |
| Approach: | They propose a method for generating synthetic question answering corpora by combining question generation and answer extraction models and filtering the results to ensure roundtrip consistency. |
| Outcome: | The proposed model achieves exact match and F1 at less than 0.1% and 0.4% from human performance on SQuAD2 and NQ. |
Copied to clipboard
| Challenge: | Existing question-answering systems are limited in their ability to test reasoning and comprehension. |
| Approach: | They propose a method to automatically extract implications from QA datasets to evaluate models' consistency . they use a heuristic to generate such questions and retrain models with implication-augmented data . |
| Outcome: | The proposed method shows that generated implications are well formed and valid . retraining with implication-augmented data improves consistency on both synthetic and human-generated implications. |
Copied to clipboard
| Challenge: | Comp-Comp is an iterative benchmarking framework grounded in the principles of comprehensiveness and compactness. |
| Approach: | They propose a benchmark framework that incorporates the principle of comprehensiveness and compactness. |
| Outcome: | The proposed framework is domain-agnostic and adaptable to a wide range of specialized fields. |
Copied to clipboard
| Challenge: | Temporal question answering is an established method for evaluating temporal reasoning in large language models. |
| Approach: | They propose a numerical estimation task where all questions require a numeric, temporal answer, allowing us to evaluate models beyond EM. |
| Outcome: | The proposed model responses are based on a numerical estimation task and are distilled from Test of Time and TempTabQA. |
Copied to clipboard
| Challenge: | Using simulated feedback, our system (called TeachMe) continually improves with time, and without model retraining. |
| Approach: | They propose to augment a QA model with a dynamic memory of user feedback, containing user-supplied corrections toerroneous model beliefs that users identify during interaction. |
| Outcome: | The proposed system improves with time and without model retraining, and with real users, by 15% on a hidden test set after teaching. |
Copied to clipboard
| Challenge: | Existing KBQA datasets are outdated and inefficient in human labor, and assisting tools like Large Language Models (LLM) are not utilized to reduce the workload. |
| Approach: | They propose a semi-automated question answering task that uses structured knowledge graphs to answer extensive knowledge-intensive questions. |
| Outcome: | The proposed approach includes KBQA, MRC, and Information Retrieval tasks for low-resource languages. |
Copied to clipboard
| Challenge: | Question answering (QA) models have shown rapid progress enabled by the availability of large, high-quality benchmark datasets. |
| Approach: | They present a multi-way aligned extractive QA evaluation benchmark in 7 languages . they evaluate state-of-the-art cross-lingual models and machine-translation-based baselines . |
| Outcome: | The proposed model is based on MLQA, which has over 12K instances in english and 5K in each other language. |
Copied to clipboard
| Challenge: | Community Question Answering (CQA) forums provide answers to many real-life questions. |
| Approach: | They propose to make Persian dataset PerCQA public to encourage more research in Persian CQA. |
| Outcome: | The proposed dataset contains 989 questions and 21,915 annotated answers from the most well-known Persian forum. |
Copied to clipboard
| Challenge: | Existing methods to extend context length of Large Language Models (LLMs) still struggle with retrieval and reasoning in long context inputs. |
| Approach: | They propose a coarse-to-fine method to enhance multi-document question-answering capacities by removing background and distracting documents. |
| Outcome: | Experiments show that CAFE outperforms baseline methods on multiple documents. |
Copied to clipboard
| Challenge: | Question-Answering (QA) has seen significant advances in recent years, achieving near human-level performance over some benchmarks. |
| Approach: | They propose to use a native QA dataset for an East African language, Tigrinya, to build similar resources for related languages. |
| Outcome: | The proposed method is applicable to constructing similar resources for related languages. |
Copied to clipboard
| Challenge: | Existing reputation systems do not take linguistic quality into account in reputation scores estimation. |
| Approach: | They build statistical models that learn reputation from syntactic and semantic structures extracted from their associated answers content. |
| Outcome: | The proposed models show that users’ writing styles play important roles in building reputation points. |
Copied to clipboard
| Challenge: | a question answering dataset is a competition that has a leaderboard that determines the best answers. |
| Approach: | They propose to apply the best practices of trivia tournaments to question answering datasets . they outline key lessons that can transfer to QA research . |
| Outcome: | The proposed model is based on the best practices of trivia tournaments . the model is used to identify the best question answering teams . |
Copied to clipboard
| Challenge: | Existing approaches to generate narrative-driven recommendation are based on large language models (LLMs) but the RAG paradigm is inherently ill-suited for such special queries. |
| Approach: | They propose a novel retrieve-rank paradigm that generatively retrieves structurally adaptive and semantically aligned candidates, ensuring both extensive candidate coverage and high-quality information. |
| Outcome: | The proposed paradigm outperforms the existing paradigm and the existing one under real-world scenarios. |
Copied to clipboard
| Challenge: | Question answering (QA) is an intuitive means to query text data. |
| Approach: | They propose a radiology question-answer-evidence-pair dataset with 3074 questions posed against radiology reports and annotated with their corresponding answer spans by physicians. |
| Outcome: | The proposed dataset has 3074 questions posed against radiology reports and annotated with their corresponding answer spans by physicians. |
Copied to clipboard
| Challenge: | Policy compliance detection is the task of ensuring that a scenario conforms to a policy. |
| Approach: | They propose to decompose policy compliance detection into question answering . they propose to use an existing dataset to augment expert annotations . |
| Outcome: | The proposed approach improves accuracy in cross-policy setups, especially when policies are unseen in training. |
Copied to clipboard
| Challenge: | a problem of information sparsity in QA tasks is causing fragmentation of textual data . highlighting entity-AWare Knowledge (HAWK) framework can be used to address this problem . |
| Approach: | a framework is proposed to highlight key information in a context and structuralize it in an entity-aware manner. |
| Outcome: | a proposed framework improves QA tasks with long contexts by highlighting key information in a context . the framework achieves a 27.6-point F1 score increase and an average win rate of 76.75% . |
Copied to clipboard
| Challenge: | Existing methods for QA data generation are limited by the dependence of existing evaluation metrics on ground truth labels. |
| Approach: | They propose a set of unsupervised evaluation metrics for QA data that enable multidimensional assessment based on the relationships among context,question and answer. |
| Outcome: | The proposed method outperforms state-of-the-art methods on public datasets and shows that it produces high-quality and domain-specific QA pairs. |
Copied to clipboard
| Challenge: | Existing question answering systems focus on extracting answers from single spans, but real-world scenarios require synthesizing information from multiple spans. |
| Approach: | They propose a dataset that leverages the MASH-QA dataset and large language models (LLMs) to ensure that each Q/A pair requires considering all selected spans. |
| Outcome: | The proposed method enables the model to answer multiple Q/A pairs in a single span, while ensuring that all selected spans are considered. |
Copied to clipboard
| Challenge: | Question Answering (QA) tasks require a mix of relevant and irrelevant information in these contexts to perform well. |
| Approach: | They propose a context filtering approach that removes non-essential details, summarizing crucial content through Reward Modeling. |
| Outcome: | The proposed approach outperforms baseline models in 6.8-folds. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable language generation capabilities, propelling advancements in various understanding/generation tasks, including opendomain question answering (QA). |
| Approach: | They propose a chain-of- Discussion framework to leverage synergy among multiple open-source Large Language Models (LLMs) aiming to provide more correct and more comprehensive answers for open-ended QA, although they are not strong enough individually. |
| Outcome: | The proposed framework leverages the synergy among multiple open-source Large Language Models (LLMs) to provide more correct and comprehensive answers for open-ended QA, although they are not strong enough individually. |
Copied to clipboard
| Challenge: | Using a dataset of tweets and Reddit, we investigate the public opinion on cryptocurrency and bitcoin on Twitter and RedDit. |
| Approach: | They create a dataset to investigate the public opinion on cryptocurrency and bitcoin on Twitter and Reddit. |
| Outcome: | The proposed dataset contains gold standard and silver standard labels and a question-answering sub-corpus. |
Copied to clipboard
| Challenge: | despite near-perfect results, effectiveness of model editing in real-world applications remains unclear. |
| Approach: | They propose QAEdit and WILD to better reflect real-world use of model editing . they propose a benchmark aligned with widely used question answering datasets and a task-agnostic evaluation framework . |
| Outcome: | The proposed QAEdit benchmark and WILD evaluation framework show that current models perform worse than previously reported. |
Copied to clipboard
| Challenge: | Open-domain question answering uses evidence retrieved from large corpus to answer questions . state-of-the-art approaches require intermediate evidence annotations for training . however, such intermediate annotations are expensive and methods that rely on them cannot transfer to the more common setting . |
| Approach: | They propose an open-domain question answering approach that alternately finds evidence from an up-to-date model and encourages the model to learn the most likely evidence. |
| Outcome: | The proposed approach improves over weak retrievers on multi-hop and single-hop benchmarks without using evidence labels. |
Copied to clipboard
| Challenge: | Large language models struggle to utilize long contexts efficiently, resulting in a question answering problem. |
| Approach: | They propose a method to generate a short document that contains the most relevant parts for a given context window. |
| Outcome: | The proposed method improves the QA task by providing a short and focused VDoc to the LLM while keeping the context window full. |
Copied to clipboard
| Challenge: | Existing frameworks for QA datasets lack regional specificity and cultural specificity. |
| Approach: | They propose a framework to quench native language QA datasets in native languages for LLM evaluation and tuning. |
| Outcome: | The proposed framework is scalable, language-independent and can be used to build culturally and regionally aligned QA datasets in native languages. |
Copied to clipboard
| Challenge: | Question-answering (QA) data often encodes essential information in many facets . a growing interest of QA has led to many large-scale QA datasets available to the community . |
| Approach: | They propose a question-answer driven sentence encoding framework to learn representations from QA data. |
| Outcome: | The proposed framework learns representations from QA data, using BERT or other state-of-the-art contextual language models. |
Copied to clipboard
| Challenge: | Existing work on Temporal Question Answering (TQA) has focused on questions anchored to specific timestamps or events. |
| Approach: | They introduce a benchmark to address present-anchored temporal QA (PATQA) which includes single and multi-hop temporal questions. |
| Outcome: | The proposed model can be automatically refreshed by re-running SPARQL queries on a knowledge graph. |
Copied to clipboard
| Challenge: | Existing approaches to consolidate textual inputs are difficult to implement . a recent study aims to capture content overlap by combining multiple textual elements . |
| Approach: | They propose to align predicate-argument relations across texts to represent content overlap . their setting exploits QA-SRL, utilizing question-answer pairs to capture predicates . |
| Outcome: | The proposed task captures content overlap beyond lexical similarity and complements cross-document coreference with proposition-level links, offering potential use for downstream tasks. |
Copied to clipboard
| Challenge: | Existing systems for text comprehension are inadequate for more holistic comprehension of a discourse. |
| Approach: | They propose a new paradigm that captures both discourse and semantic links between sentences in the form of free-form, open-ended questions. |
| Outcome: | The proposed model captures discourse and semantic links between sentences in the form of free-form, open-ended questions. |
Copied to clipboard
| Challenge: | Visual question answering (VQA) is a task that requires an understanding of both the image and the question to provide a natural language answer. |
| Approach: | They propose a multimodal framework that leverages language guidance to answer questions more accurately. |
| Outcome: | The proposed framework improves on the multi-choice question-answering task using CLIP and BLIP models. |
Copied to clipboard
| Challenge: | OpenCQA is a task to answer open-ended questions about charts with descriptive texts. |
| Approach: | They propose a task to answer open-ended questions about charts with descriptive texts. |
| Outcome: | The proposed task is to answer an open-ended question about a chart with descriptive texts. |
Copied to clipboard
| Challenge: | Existing data synthesis methods generate simplistic and homogeneous QA pairs with limited scale and diversity. |
| Approach: | They propose a framework to synthesize large-scale, diverse, and high-quality QA data for mid-training. |
| Outcome: | The proposed framework improves on 500B-token BoostQA data over pre-training benchmarks. |
Copied to clipboard
| Challenge: | Existing XQA methods focus on reasoning on a single knowledge source, e.g., structured knowledge bases, unstructured corpora, etc. Existing work in XQA focuses on integrating information from heterogeneous knowledge sources. |
| Approach: | They propose to leverage question decomposing for heterogeneous knowledge integration by breaking down a complex question into simpler ones and selecting the appropriate knowledge source for each sub-question. |
| Outcome: | The proposed framework outperforms SOTA methods on complex QA datasets. |
Copied to clipboard
| Challenge: | graph neural networks capture structured graph information, but lack integration at the reasoning level. |
| Approach: | They propose a framework that leverages graph structural information to reason interpretable academic QA results. |
| Outcome: | The proposed framework outperforms sota baselines on OpenAlex and DBLP datasets. |
Copied to clipboard
| Challenge: | Knowledge graph question answering (KGQA) aims to answer natural language questions using knowledge graphs. |
| Approach: | They propose a framework that retrieves refined reasoning paths and evaluates their sufficiency. |
| Outcome: | The proposed framework outperforms existing baselines while enabling small open-source LLMs to achieve competitive results without fine-tuning LLM. |
Copied to clipboard
| Challenge: | Temporal misalignment is a problem for knowledge-intensive tasks where models must rely on data from the past to make predictions. |
| Approach: | They propose a temporal misalignment task to predict how long a given fact will remain true. |
| Outcome: | The proposed task improves calibration for knowledge-intensive tasks under temporal misalignment by discarding volatile facts. |
Copied to clipboard
| Challenge: | Pretrained sequence-to-sequence (seq2sequ) models have been widely used to solve extractive tasks, where parts of the input are extracted to form the desired output. |
| Approach: | They propose a simple fix to tokenization inconsistency that damages extractive nature of generative models by causing performance drop and hallucination. |
| Outcome: | The proposed model performs better in both in-domain and out-of-domain datasets with a notable average of +1.7 F1 gain when a BART model is trained on SQuAD and evaluated on 8 QA datasets. |
Copied to clipboard
| Challenge: | Question and answer generation (QAG) is a task of generating question-answer pairs given a context. |
| Approach: | They propose to leverage sequence-to-sequence language model fine-tuning to generate question-answer pairs given a context. |
| Outcome: | The proposed model outperforms other more convoluted approaches in the end-to-end model and is computationally light at both training and inference times. |
Copied to clipboard
| Challenge: | Question answering (QA) is a fundamental task in the field of Natural Language Processing (NLP). |
| Approach: | They propose a database querying and reasoning dataset for question answering that is designed to accommodate sequential questions and multi-hop queries. |
| Outcome: | The proposed dataset better mirrors the dynamics of real-world information retrieval and analysis with a particular focus on the financial reports of US companies. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have been successful in understanding language and processing text, but their cost prohibits their practical applications. |
| Approach: | They propose a multi-agent collaboration method that breaks down lengthy documents into smaller, more manageable chunks and organizes the member agents to read their assigned chunks. |
| Outcome: | The proposed method achieves 16.42% and 1.63% accuracy gains over existing models on single-hop and multi-hop QA settings. |
Copied to clipboard
| Challenge: | Existing studies on the effectiveness of moral self-correction in large language models have not been conducted. |
| Approach: | They propose that moral self-correction is a computationally efficient method for reducing harmful content in LLMs. |
| Outcome: | The proposed method reduces harmful content in LLMs, but it remains under-explored . it can help LLM find shortcut to more morally correct output, the authors argue . |
Copied to clipboard
| Challenge: | Iterative retrieval-augmented generation models are difficult to use for multihop question answering (QA) . their retrieval processes can be disrupted by irrelevant documents or factually inaccurate chain-of-thoughts . |
| Approach: | They propose a knowledge-driven iterative retriever model that decomposes documents into knowledge triples and performs iterativ retrieval with these triples to enable a factually reliable retrieval process. |
| Outcome: | The proposed model outperforms existing iRAG models with an average improvement of 9.40% in R@3 and 5.14% in F1 on multi-hop QA datasets. |
Copied to clipboard
| Challenge: | Existing work shows that large language models generate incorrect statements due to over-reliance on parametric knowledge. |
| Approach: | They propose a framework that utilizes syntax trees to guide information retrieval and reasoning for question answering. |
| Outcome: | The proposed framework improves on existing state-of-the-art methods for large-scale query processing. |
Copied to clipboard
| Challenge: | Existing work on citation generation has focused on unambiguous settings with single answers, failing to address the complexity of real-world scenarios. |
| Approach: | They propose a task of QA with source citation in ambiguous settings where multiple valid answers exist, where multiple sources exist. |
| Outcome: | The proposed framework generates multiple answers and cites their sources, allowing users to verify the factuality of each answer and make informed decisions. |
Copied to clipboard
| Challenge: | Recent studies have highlighted the significant performance variation that can arise from minor changes in prompt design. |
| Approach: | They propose to tokenize the space following the colon to facilitate automated answer extraction via next-token probabilities. |
| Outcome: | The proposed tokenization improves model calibration and improves confidence estimates. |
Copied to clipboard
| Challenge: | Existing methods to detect and safeguard LLMs against knowledge leakage fail to address the long-term challenge of mitigating it. |
| Approach: | They propose a method to reinforce and safeguard existing benchmarks against knowledge leakage by perturbation-based detection and counterfactual rewriting to disrupt memorization while preserving original intent. |
| Outcome: | The proposed method reduces memorization effects in long-context QA benchmarks, providing a more accurate assessment of model reasoning and generalization abilities. |
Copied to clipboard
| Challenge: | Our Dataset is the first cross-lingual QA dataset with a focus on African languages. |
| Approach: | They propose to use African languages as the only high-coverage source of answer content for cross-lingual open-retrieval question answering systems. |
| Outcome: | Our Dataset includes 12,000+ XOR QA examples across 10 African languages. |
Copied to clipboard
| Challenge: | Historical newspapers from the colonial period offer valuable evidence of how racializing language evolved over time. |
| Approach: | They propose a contextual question answering and visual question answering task from colonial newspapers . they propose linguistic training for temporal word embedding with a compass to study racialization . |
| Outcome: | The proposed tasks are limited for low-resource tasks, the authors show . the authors compare the results of two QA pairs from colonial newspapers to a compass . |
Copied to clipboard
| Challenge: | Multiple-choice exam questions with “None of the above” (NA) options have been extensively studied in educational testing . however, their impact on Large Language Models (LLMs) evaluation remains underexplored . |
| Approach: | They conduct systematic experiments with 28 LLMs on the MMLU benchmark to examine how NA options affect model performance and confidence calibration. |
| Outcome: | The results highlight important implications for benchmark design and raise questions about LLMs’ ability to handle uncertainty in real-world applications. |
Copied to clipboard
| Challenge: | Long-form question answering (LFQA) answers are prone to hallucinations and factual inconsistencies, challenging their faithful evaluation. |
| Approach: | They propose a dataset with localized error annotations for human-written and model-generated LFQA answers. |
| Outcome: | The proposed approach reduces errors and improves quality of the answers across multiple models. |
Copied to clipboard
| Challenge: | Large language models (LLMs) possess strong capabilities in language understanding and generation, as well as remarkable problem-solving abilities. |
| Approach: | They propose a benchmark to assess the cognitive alignment capabilities of large language models in educational QA. |
| Outcome: | The proposed evaluation benchmark assesses the cognitive alignment capabilities of large language models in educational QA. |
Copied to clipboard
| Challenge: | Traditional pre-trained LLMs struggle with domain-specific terminology, while fine-tuned LLM requires substantial computational resources. |
| Approach: | They propose a training-free approach that combines TF-IDF with prompt-based LLMs to address technical questions. |
| Outcome: | The proposed system improves the accuracy and efficiency of QA systems in technical domains without LLM retraining. |
Copied to clipboard
| Challenge: | Recent work in NLP has taken advantage of question generation capabilities of LLMs to enhance a wide range of applications. |
| Approach: | They propose a salience predictor for inquisitive questions that is instruction-tuned . they show that highly salient questions are empirically more likely to be answered in the same article . |
| Outcome: | The proposed model is based on linguist-annotated salience scores of 1,766 questions . it shows that answering salient questions improves comprehension of the text . |
Copied to clipboard
| Challenge: | Multilingual question answering systems must ensure factual consistency across languages while also accounting for cultural variation in subjective responses. |
| Approach: | They propose a user-in-the-loop fact-checking pipeline to detect factual and cultural discrepancies in multilingual QA knowledge bases. |
| Outcome: | The proposed tool detects factual and cultural discrepancies in bilingual question answering systems. |
Copied to clipboard
| Challenge: | Existing question-answering datasets are expensive and difficult to annotate and time-consuming to gather. |
| Approach: | They propose to transform Manchester questions into web queries using the same question datasets. |
| Outcome: | The proposed questions can be trained on a Manchester QA dataset using the Quiz Bowl (QB) sample. |
Copied to clipboard
| Challenge: | Multiple-choice question answering (MCQA) is easy to evaluate but adds a meta-task . prior work has shown that language models exhibit selection biases for particular option identifiers such as the label "A" |
| Approach: | They find that option-boundary residual states contain strong linearly decodable signals . winning content position becomes decoded after final option is processed . |
| Outcome: | The proposed model solves the problem and outputs the symbol that represents the answer. |
Copied to clipboard
| Challenge: | Temporal relation annotation in the clinical domain is crucial but challenging due to its workload and the medical expertise required. |
| Approach: | They propose an annotation method that integrates event start-points ordering and question-answering as the annotation format. |
| Outcome: | The proposed method achieves a 0.72 F1 score and enables collaboration among medical experts and non-experts. |
Copied to clipboard
| Challenge: | Retrieval-augmented generation grounds language models in external evidence, but multi-hop question answering remains difficult . iterative pipelines must control what to retrieve next and when evidence is adequate. |
| Approach: | They propose an iterative framework with an explicit controller, S2G-Judge . they map structured gap items into the next retrieval query to produce stable retrieval trajectories . |
| Outcome: | Experiments on TriviaQA, HotpotQA, and 2WikiMultiHopQA show that S2G-RAG improves multi-hop QA performance and robustness under multi-turn retrieval. |
Copied to clipboard
| Challenge: | Existing methods to condense extensive documents with no loss of information are difficult to implement in real-world scenarios. |
| Approach: | They propose a framework that employs an active strategy to condense extensive documents without losing key information. |
| Outcome: | The proposed framework improves performance and compression rate on multi-hop question-answering benchmarks. |
Copied to clipboard
| Challenge: | Recent advances in large language models have led to claims of AI surpassing humans in QA tasks . authors: models are purportedly acing tests that many humans find challenging . |
| Approach: | They propose a framework that enables quantitative assessment and comparison of problem-solving abilities in QA agents. |
| Outcome: | The proposed framework uncovers distinctficiency patterns in knowledge domains and reasoning skills. |
Copied to clipboard
| Challenge: | Log files are crucial for monitoring, diagnostics, and root cause analysis in IT systems . their sheer volume makes manual analysis overwhelming and traditional methods are ineffective . |
| Approach: | They propose a framework that constructs a multi-entity temporal hypergraph using log attribute-value pairs as nodes and connects them with hyperedges. |
| Outcome: | The proposed framework is model-agnostic and training-free and scales with open-source LLMs. |
Copied to clipboard
| Challenge: | Recent literature reveals that supervised fine-tuning (SFT) is suboptimal for domain-specific question-answering tasks. |
| Approach: | They propose a query diversification strategy for robust conflict detection and a knowledge-aware fine-tuning approach to effectively boost LLMs’ performance. |
| Outcome: | The proposed approach improves the model generalization and alleviates the hallucination. |
Copied to clipboard
| Challenge: | Existing reranking frameworks optimize semantic relevance, leading to unstable rankings and opaque decisions on long documents. |
| Approach: | They propose a structured reranking framework that reframes financial evidence selection as constraint satisfaction under a finance-aware schema. |
| Outcome: | FINCARDS improves early-rank retrieval over lexical and LLM-based reranking baselines while reducing ranking variance. |
Copied to clipboard
| Challenge: | Limited availability of multilingual text corpora for pretraining results in poor performance on downstream tasks due to undertrained representation spaces for languages other than English. |
| Approach: | They propose a method that integrates source and target language representations within low-rank (LoRA) adapters using lightweight linear transformations to enhance representation quality and transfer performance for languages other than English. |
| Outcome: | The proposed method improves representation quality and performance for languages other than English while maintaining parameter efficiency. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have achieved remarkable success in natural language processing (NLP), particularly in single-turn question answering (QA) on short-text. |
| Approach: | They propose a framework that captures logical correlations across chunks of ELC and maintains coherence of multi-turn Questions. |
| Outcome: | The proposed framework is able to capture logical correlations across chunks of ELC and maintain coherence of multi-turn Questions. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are typically trained on static datasets, preventing them from integrating real-time updates. |
| Approach: | They propose a dynamic question-answer answering dataset reflecting real-world knowledge updates that are automatically compared between Wikipedia versions and generating question-anchor pairs based on these updates. |
| Outcome: | The proposed framework improves LLMs' performance on time-sensitive question answering by maintaining a dynamic knowledge updating process. |
Copied to clipboard
| Challenge: | Existing annotated datasets for NLP tasks in languages with limited resources are limited. |
| Approach: | They propose to use machine translation to convert existing Tigrinya dataset into a Tigrina dataset in SQuAD format. |
| Outcome: | The proposed dataset is an expert-annotated Tigrinya dataset with 2,685 question-answer pairs covering 122 diverse topics. |
Copied to clipboard
| Challenge: | Existing Question Answering systems are limited by noisy documents and flawed QA pairs. |
| Approach: | They propose a high-quality subset of NarrativeQA focused on literary works . they identify and correct low-quality QA samples while removing extraneous text . |
| Outcome: | The proposed subset of NarrativeQA is based on literary works. |
Copied to clipboard
| Challenge: | Large language models are promising for medical question answering in china, but remain unreliable due to hallucinations, weak factual grounding and difficulty handling clinically complex cases. |
| Approach: | They propose a framework that combines hierarchical medical adaptation with complexity-aware expert routing for reliable Chinese medical QA. |
| Outcome: | The proposed framework outperforms strong general and medical LLM baselines on four Chinese medical benchmarks. |