Papers with question-answering
Copied to clipboard
| Challenge: | Text adventure games provide a stepping stone toward grounding action in language . prior work demonstrated that using a knowledge graph as a state representation facilitates faster control policy learning. |
| Approach: | They propose to use knowledge graphs as a representation for domain knowledge transfer for training text-adventure playing reinforcement learning agents. |
| Outcome: | The proposed methods let us learn a higher-quality control policy faster in text adventure games. |
Copied to clipboard
| Challenge: | Using Amazon reviews, we find that the answer to a question is only in 45% of cases. |
| Approach: | They combine Amazon reviews with consumer reviews and manually analyse 400 questions from four domains to find that reviews directly contain the answer to the question . they then compare QA systems that use reviews in addition to the questions to see if they can be useful for other question types. |
| Outcome: | The proposed system outperforms the chance baseline but not by a large margin. |
Copied to clipboard
| Challenge: | Recent research shows that pretrained language models are often brittle for complex reasoning tasks. |
| Approach: | They propose to use pre-trained language models to teach machines to reason over texts . they will review recent promising approaches to tackling complex reasoning tasks . |
| Outcome: | This tutorial reviews promising approaches to complex reasoning tasks . it reviews the methods that can be used to augment models with robustness . |
Copied to clipboard
| Challenge: | Despite the proven efficacy of deep neural networks at-large, their opaqueness is a major cause of concern. |
| Approach: | They will present research work on interpreting fine-grained components of a neural network model from two perspectives, i) fine-grain interpretation, and ii) causation analysis. |
| Outcome: | This paper presents work on interpreting fine-grained components of a neural network model from two perspectives, i) fine-grain interpretation, and ii) causation analysis. |
Copied to clipboard
| Challenge: | Existing libraries for building R-LLMs provide high-level abstractions without sufficient transparency for evaluating and optimizing prompts within specific inference processes. |
| Approach: | They propose an open-source framework to facilitate the development, evaluation, and optimization of R-LLMs for knowledge-intensive tasks. |
| Outcome: | The framework improves hand-crafted prompts, inference processes and quantitatively measures overall system performance. |
Copied to clipboard
| Challenge: | Recent advances in AI can be attributed to the remarkable performance of Large Language Models (LLMs) success of LLMs depends on specific training techniques, such as instruction tuning and prompting . |
| Approach: | They explore the capabilities of Large Language Models (LLMs) in various tasks and languages . they also examine their performance, fine-tuning, instructions tuning, and close vs. open models . |
| Outcome: | The proposed model can be used for speech and multimodal tasks across modalities, languages, and dialects. |
Copied to clipboard
| Challenge: | despite considerable progress, most machine reading comprehension tasks lack sufficient training data to fully exploit powerful deep neural network models. |
| Approach: | They propose to use QA data to generate more training data for machine reading comprehension tasks by crowdsourcing . they first collect a large-scale multiple-choice QA dataset for Chinese, ExamQA, and then use incomplete, yet relevant snippets returned by a web search engine as the context for each QA instance. |
| Outcome: | The proposed model improves a Chinese MRC task with +5.1% accuracy and +3.8% exact match. |
Copied to clipboard
| Challenge: | Recent Deep Learning (DL) models have achieved human-level accuracy on natural language tasks such as question-answering, natural language inference, and textual entailment. |
| Approach: | They propose an unsupervised question-answering based approach for a similar task, fact-checking. |
| Outcome: | The proposed approach achieves label accuracy of 80.2% on the development set and 80.25% on the test set. |
Copied to clipboard
| Challenge: | Large language models (LLMs) integrated with retrieval-augmented generation (RAG) are a dominant framework for building intelligent assistants. |
| Approach: | They propose a benchmark to evaluate LLMs' reasoning capability over real-world conflicting documents retrieved from the web. |
| Outcome: | The proposed benchmark evaluates LLMs' reasoning capability over real-world conflicting documents retrieved from the web. |
Copied to clipboard
| Challenge: | Traditional visual question generation (VQG) focuses on single images, resulting in a limited ability to comprehend time-series information of the underlying event. |
| Approach: | They propose to generate engaging questions from multiple images using a visual question generation dataset and establish a series of baselines. |
| Outcome: | The proposed model builds stories behind the image sequence to allow for creativity and experience sharing and hence draw attention to downstream applications. |
Copied to clipboard
| Challenge: | Knowledge graphs are incomplete in the information they represent, necessitating knowledge graph completion tasks. |
| Approach: | They propose a new entity/relation embedding layer that learns to differentiate distinctive entity and relation types, thus allowing the model to learn the structure of the knowledge graph. |
| Outcome: | The proposed language model learns to differentiate distinct entity and relation types, thus learning the structure of the knowledge graph. |
Copied to clipboard
| Challenge: | Neural approaches have improved machine comprehension tasks, but models often operate as a black-box, resulting in lower interpretability. |
| Approach: | They propose a hybrid approach to quantify model uncertainty using Bayesian weight approximation and boost up inference speed by 80% relative to test time. |
| Outcome: | The proposed approach boosts inference speed by 80% relative to the previous approach and is applied to a clinical dialogue comprehension task. |
Copied to clipboard
| Challenge: | Recent years have witnessed an explosion of Large Language Models (LLMs), with impressive performance on various NLP tasks. |
| Approach: | They propose to use image-based representations to compare LLMs' performance on table-related tasks such as question-answering and fact-checking to determine their effectiveness. |
| Outcome: | The proposed model performs better on image-based representations than on text-based models. |
Copied to clipboard
| Challenge: | Recent advances in foundation language models have shown the efficacy of pre-trained models across diverse QA tasks. |
| Approach: | They propose a multi-task benchmark for evaluating causality-aware language models to unify causal QA research. |
| Outcome: | The proposed model outperforms single-task fine-tuned models on the CALM-Bench tasks. |
Copied to clipboard
| Challenge: | Named entity linking (NEL) is a preprocessing step in commercial systems . a small organization or individual could use an off-the-shelf system to accomplish the same objectives . |
| Approach: | They examine how to repurpose off-the-shelf NEL systems to correct sport-related errors. |
| Outcome: | The proposed model can improve sports question-answering accuracy by 25% . the proposed model is based on the best available model . |
Copied to clipboard
| Challenge: | Retrieval-augmented generation (RAG) is a key framework in natural language processing . however, the effectiveness of RAG is often hindered by coreferential complexity in retrieved documents . |
| Approach: | They investigate how entity coreference affects document retrieval and generative performance in RAG-based systems. |
| Outcome: | The proposed model improves QA performance and retrieval relevance and contextual understanding. |
Copied to clipboard
| Challenge: | Currently, the evaluation of large language models (LLMs) such as ChatGPT in academic datasets is difficult due to the difficulty of evaluating the generative outputs produced by this model against the ground truth. |
| Approach: | They evaluate ChatGPT across 140 tasks and analyze 255K responses it generates in academic datasets. |
| Outcome: | The proposed model performs well on 140 tasks and generates 255K responses in these datasets. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a foundational task for a variety of applications like question answering and machine translation. |
| Approach: | They propose to split entity recognition problem into two sub-tasks and optimize them separately for each sub-task. |
| Outcome: | The proposed system outperforms baselines on OntoNotes5.0, WNUT17 and a cybersecurity dataset and gives on-par performance on BioNLP13CG. |
Copied to clipboard
| Challenge: | Existing methods for probability-based UE are limited by their inability to handle biased probabilities and complex semantic dependencies between tokens. |
| Approach: | They propose a learning-based scoring function that captures complex dependencies between tokens and probabilities and produces more reliable responses. |
| Outcome: | The proposed function outperforms existing scoring functions in question-answering and arithmetical reasoning tasks with different datasets. |
Copied to clipboard
| Challenge: | Existing methods for generating summarizations using QA-based supervision produce higher quality summaries than baseline methods. |
| Approach: | They propose a method for incorporating question-answering signals into a summarization model by automatically marking document NPs as salient based on whether they are answered in the gold summaries. |
| Outcome: | The proposed method generates higher-quality summaries than baseline methods on benchmark summarization datasets. |
Copied to clipboard
| Challenge: | Existing retrieval systems assume relevant corpora are fully (e.g., publicly) accessible, but users are often unwilling to expose their private data to entities hosting public data. |
| Approach: | They propose a split iterative retrieval problem involving iterating retrieval over multiple privacy scopes and propose 'concurrentQA' benchmark to test this problem. |
| Outcome: | The proposed method improves on the existing retrieval methods but still suffers performance degradations when applied to a dataset from a public and private distribution. |
Copied to clipboard
| Challenge: | Existing studies have focused on disfluency detection and removal, with limited studies into its impact on downstream tasks. |
| Approach: | They propose to incorporate disfluency in summarization models to reduce the impact of replacement disfluencies on natural language processing tasks. |
| Outcome: | The proposed model improves on both public and real-life datasets and shows that it can handle disfluent data with up to 6.99-point degradation in Rouge-L score and replacement disfluencies have the highest negative impact. |
Copied to clipboard
| Challenge: | Large language models (LLMs) excel in various tasks, but often produce hallucinations . retrieved contexts, misrepresent information, or generate outright contradictions . |
| Approach: | They propose a framework that measures hallucination faithfulness of large language models . they introduce a leaderboard that leverages diverse human-annotated hallucinian examples . |
| Outcome: | The proposed framework improves hallucination evaluations by leveraging human-annotated examples. |
Copied to clipboard
| Challenge: | Existing datasets for human-like dialogue tasks are deficient due to the complexity of human conversations. |
| Approach: | They construct a large-scale Chinese E-commerce conversation corpus with 1 million dialogues, 20 million utterances, and 150 million words. |
| Outcome: | The proposed dataset includes 1 million multi-turn dialogues, 20 million utterances, and 150 million words. |
Copied to clipboard
| Challenge: | Hallucination is a well-known phenomenon in text generated by large language models . state-of-the-art LLMs still have a number of weaknesses, including the tendency to generate hallucinatory statements without considering the factuality . |
| Approach: | They propose a dataset that captures hallucinations made by retrieval-augmented LLMs . they propose to use these methods to help detect hallucinosity in QA tasks . |
| Outcome: | The proposed method captures hallucinations made by retrieval-augmented LLMs for QA tasks. |
Copied to clipboard
| Challenge: | Existing AMR metrics are inefficient and struggle to capture semantic similarity . Existing metrics are not efficient and lack a systematic evaluation benchmark . |
| Approach: | They propose a new AMR similarity metric, rematch, which matches graphs structurally and semantically to each other. |
| Outcome: | The proposed metric is five times faster than the next most efficient metric. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are prone to hallucination and rely on static, pre-annotated references for evaluation. |
| Approach: | They propose a framework to assess large language models without fixed ground-truth answers by iteratively generating web queries and synthesizing external evidence. |
| Outcome: | The proposed framework achieves substantial to perfect agreement with human evaluations on multiple free-form QA benchmarks. |
Copied to clipboard
| Challenge: | a new generation of English-oriented Large Language Models significantly outperforms older LLMs on low-resource languages. |
| Approach: | They compare Bengali-oriented LLMs with open-weight and closed-source LLM models . they conclude that there is a need for a Bengali model, but lacks high-quality pretraining data . |
| Outcome: | The proposed model outperforms existing models on Bengali on low-resource languages . the results highlight biases in machine-translated datasets used for Bengali NLP tasks . |
Copied to clipboard
| Challenge: | Recent attempts at screenplay summarization focus on fine-tuning transformer-based pre-trained models, but these models often fall short in capturing long-term dependencies and latent relationships. |
| Approach: | They propose a novel resource that represents movie scripts as a movie character-aware discourse graph (CaD Graph) this resource aims to preserve all salient information, offering a more comprehensive and faithful representation of the screenplay’s content. |
| Outcome: | The proposed model preserves all salient information, offering a more comprehensive and faithful representation of the screenplay’s content. |
Copied to clipboard
| Challenge: | Cognitive science has long promoted the formation of mental models as central to understanding and question-answering. |
| Approach: | They train a new model, DREAM, to answer questions that elaborate the scenes that situated questions are about and then provide those elaborations as additional context to a question-answering (QA) model. |
| Outcome: | The proposed model is able to create better scene elaborations than a representative state-of-the-art, zero-shot model. |
Copied to clipboard
| Challenge: | Existing methods for QA are hampered by increased training costs . current methods suffer significant performance degradation when applied to out-of-domain examples. |
| Approach: | They propose a method that combines prompting methods and linear probing with fine-tuning strategy, which does not entail additional cost. |
| Outcome: | The proposed method outperforms state-of-the-art baselines with an average increase in F1 score of 4.5%-7.9%. |
Copied to clipboard
| Challenge: | Supervised self-training methods have transformed applied machine learning . however, adapting to target data has received little attention . |
| Approach: | They propose a method to generate synthetic QA pairs for unsupervised self adaptation . they use massive amounts of data to simulate self-supervised tasks . |
| Outcome: | The proposed method improves QA systems significantly by using less data and training computation than existing augmentation approaches. |
Copied to clipboard
| Challenge: | Recent methods supervise only the final answer accuracy using reinforcement learning with verifiable rewards (RLVR). |
| Approach: | They propose to train search agents to search and reason over scientific papers and a factoid QA dataset with 60k biomedical paper abstracts. |
| Outcome: | The proposed model outperforms non-RL retrieval baselines and is scalable and extendable to other scientific domains. |
Copied to clipboard
| Challenge: | Existing work on predicate entailment detection from typed open relation triples has not been able to detect predicates. |
| Approach: | They propose a pipeline for building Chinese entailment graphs using an open relation extraction method. |
| Outcome: | The proposed pipeline outperforms monolingual and Chinese entailment graphs on a parallel dataset. |
Copied to clipboard
| Challenge: | Financial dialogue transcripts pose a unique challenge for sentence-level information extraction due to their informal structure, domain-specific vocabulary, and variable intent density. |
| Approach: | They propose a framework for extracting user intent–relevant sentences from financial service calls. |
| Outcome: | The proposed framework shows strong precision and F1 performance on real-world transcripts . financial transcripts are a challenge due to their informal structure and domain-specific vocabulary . |
Copied to clipboard
| Challenge: | Large language models struggle with input errors, often failing to interpret user intent or altering the original question’s structure (over-correction). |
| Approach: | They propose a framework that uses reinforcement learning to address misinterpretation and over-correction by integrating external knowledge with the input. |
| Outcome: | The proposed framework unlocks the full potential of LLMs for the question correction task. |
Copied to clipboard
| Challenge: | Recent introduction of robust, general-purpose models for fine-tuning has enabled improvements in general natural language understanding (NLU) but such benchmarks are only available for a handful of languages. |
| Approach: | They propose a multi-task benchmark for the Polish language understanding with an online leaderboard . they also propose GLUE, a task for named entity recognition and sentiment analysis . |
| Outcome: | The proposed model performs best on three out of nine tasks in the Polish language . the proposed model is also used in an e-commerce domain to analyze the sentiments of users . |
Copied to clipboard
| Challenge: | Large language models (LLMs) generate unreliable responses due to their cognitive alignment of context and intent. |
| Approach: | They propose a benchmark to identify possible implicit assumptions in QA questions . they use retrieved Wikipedia fragments to identify interpretations for a given query . |
| Outcome: | The proposed benchmark identifies possible implicit assumptions and improves answer accuracy by 11.75% . retrieved Wikipedia fragments help identify possible interpretations for a given query . |
Copied to clipboard
| Challenge: | Currently, video-based dialogue systems rely on a single dialogue type, hindering their versatility in practical applications. |
| Approach: | They propose to generate video-driven multilingual mixed-type dialogues using KwaiChat . they propose to create a video-based multilingual mix of 4 dialogue types, 30 domains, 4 languages, 13 topics . |
| Outcome: | The proposed model performs best on KwaiChat but is not perfect in this situation. |
Copied to clipboard
| Challenge: | Large language models have shown promise for generative and knowledge-intensive tasks including question-answering (QA) but the practical deployment still faces challenges, notably the issue of “hallucination”, where models generate plausible-sounding but unfaithful or nonsensical information. |
| Approach: | They propose a self-reflection methodology that incorporates knowledge acquisition and answer generation to address the issue of "hallucination" they use a set of LLMs to generate a more accurate and factually accurate answer. |
| Outcome: | The proposed approach improves factuality, consistency, and entailment of the generated answers. |
Copied to clipboard
| Challenge: | Existing question-answering systems focus on answering individual questions, assuming they are devoid of context. |
| Approach: | They propose to ask multiple related questions in a dataset that includes human-authored questions. |
| Outcome: | The proposed system can answer human-authored questions better than existing systems. |
Copied to clipboard
| Challenge: | a system that can show how its answers are implied by its own internal beliefs via a systematic chain of reasoning would allow better understanding of why a model produced the answer it did. |
| Approach: | They propose to combine a backward-chaining model with a verifier that checks that the model itself believes those premises through self-querying to generate multistep chains that are both faithful (the answer follows from the reasoning) |
| Outcome: | The proposed model generates chains that are faithful and truthful while maintaining answer accuracy. |
Copied to clipboard
| Challenge: | Retrieval-augmented generation (RAG) is a common technique for grounding language models in domain-specific information. |
| Approach: | They propose a new retrieval technique that incorporates diversity into the retrieval step to improve performance on reasoning-intensive QA benchmarks. |
| Outcome: | The proposed method outperforms baselines on reasoning-intensive QA benchmarks by 4–10%. |
Copied to clipboard
| Challenge: | Current datasets for conversational question answering lack realistic, domain-specific training data. |
| Approach: | They propose a model that generates question-answer representations across dialogue turns . they use flow propagation training to improve conversational flow and fluidity . |
| Outcome: | The proposed model outperforms answer-aware and answer-unaware SOTA baselines significantly . it generates different types of questions with improved fluidity and coreference alignment. |
Copied to clipboard
| Challenge: | Flowcharts are typically presented as images, driving the trend of using vision-language models for end-to-end flowchart understanding. |
| Approach: | They propose a vision-language model (VLM) that generates textual representations from flowchart images and a textual Reasoner that performs question-answering based on the text representations. |
| Outcome: | Experiments on the FlowVQA and FlowLearn benchmarks demonstrate TextFlow’s state-of-the-art performance as well as its robustness. |
Copied to clipboard
| Challenge: | Existing approaches to answer open domain questions rely on unlabeled text or synthetically generated question-answer pairs. |
| Approach: | They propose a large-scale open-domain question-answering dataset based on the Common Crawl project that can be used to in-domain pre-train popular language models. |
| Outcome: | The proposed dataset achieves promising results in zero-shot, low resource and fine-tuned settings across multiple tasks, models and benchmarks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have enabled advances in the field of natural language processing . however, their application and potential are still underexplored . |
| Approach: | They evaluate four state-of-the-art instruction-tuned Large Language Models on 13 NLP tasks in English. |
| Outcome: | The evaluated models outperform state-of-the-art models on 13 real-world clinical and biomedical NLP tasks in English. |
Copied to clipboard
| Challenge: | a lack of diverse and comprehensive question-answering datasets exists in under-resourced languages like Bangla. |
| Approach: | They propose a reading comprehension-based Bangla question-answering dataset . the dataset includes answerable and unanswerable questions covering four categories of questions . |
| Outcome: | The proposed dataset shows that it performs well as a training resource in high-resource languages. |
Copied to clipboard
| Challenge: | Existing automatic question generation datasets focus on English, resulting in data gaps for other languages. |
| Approach: | They propose a cross-lingual transfer method that allows models to generate questions in low-resource languages. |
| Outcome: | The proposed method outperforms other models and achieves comparable performance across languages. |
Copied to clipboard
| Challenge: | Extract-Refine-Retrieve-Read is a query optimization framework for large language models . it is designed to bridge the pre-retrieval information gap in Retriev-Augmented Generation systems . |
| Approach: | They propose a framework to extract parametric knowledge from Large Language Models and refine them using a specialized query optimizer. |
| Outcome: | The extract-refine-retrieve-read framework outperforms baselines on QA datasets . it is designed to meet the knowledge requirements of large language models (LLMs) |
Copied to clipboard
| Challenge: | Using publicly available materials science text data, we construct a benchmark for evaluating the performance of natural language processing (NLP) models on materials science texts. |
| Approach: | They propose a natural language benchmark for evaluating the performance of natural language processing (NLP) models on materials science text. |
| Outcome: | The proposed model outperforms BERT-based models on scientific text and a model pretrained on materials science journals. |
Copied to clipboard
| Challenge: | Existing ethical and safety considerations for large language models are important for deployment . however, some ethical concerns have been raised due to the presence of private, sensitive, or harmful information in the training data. |
| Approach: | They propose a framework that learns prompt tokens that are prepended to a query to induce unlearning in LLMs. |
| Outcome: | The proposed method improves the trade-off between utility and forgetting for text classification and question-answering. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are powerful at question-answering but prone to hallucinations due to limited domain-specific or up-to-date knowledge. |
| Approach: | They propose a framework for IDentifying RAG properties in LLM services that integrates LLMs with retrieval systems and adds an external retriever and knowledge database to mitigate hallucinations. |
| Outcome: | The proposed framework detects RAG-enhanced LLMs with 99.97% accuracy with partial or no optional knowledge and nearly 100% when the LLM and database are known. |
Copied to clipboard
| Challenge: | Recent advances in large language models have demonstrated impressive language understanding and generation capabilities, enabling them to answer a wide range of questions across various domains. |
| Approach: | They propose a refusal mechanism that instructs LLMs to refuse to answer challenging questions in order to avoid errors. |
| Outcome: | The proposed approach improves the controllability and reliability of large language models and their ability to answer questions across domains. |
Copied to clipboard
| Challenge: | False. a free-form question answering dataset can serve as a useful research benchmark for source code comprehension. |
| Approach: | They propose a free-form question answering dataset for source code comprehension . they implement syntactic rules and semantic analysis to transform code comments into question-answer pairs. |
| Outcome: | The proposed dataset can serve as a useful research benchmark for source code comprehension. |
Copied to clipboard
| Challenge: | Existing approaches to retrieve entity information are limited by document level retrieval and intermingled storage of information from different entities. |
| Approach: | They propose a framework that enhances entity-specific query handling . MES-RAG introduces proactive security measures that ensure system integrity . |
| Outcome: | Experimental results show that MES-RAG improves accuracy and recall . the framework can be integrated into existing RAG architectures . |
Copied to clipboard
| Challenge: | Existing models of geometric reasoning are based on visual representations of objects and objects, but they are not based in symbols or words. |
| Approach: | They propose a new deep network architecture that specializes in answering questions that admit latent visual representations and learns to generate and reason over such representations. |
| Outcome: | The proposed model can generate and reason over latent visual representations and is validated by two synthetic benchmarks. |
Copied to clipboard
| Challenge: | Large pre-trained language models (LMs) have a surprising ability to perform zero-shot learning. |
| Approach: | They propose to fine-tune pre-trained language models to optimize the zero-shot learning objective by aggregating 43 existing datasets and annotating 441 label descriptions in a question-answering format. |
| Outcome: | The proposed model outperforms a same-sized QA model and the previous SOTA zero-shot learning system on unseen tasks. |
Copied to clipboard
| Challenge: | Multiple-choice questions (MCQs) are widely used in the evaluation of large language models (LLMs) however, there are concerns about whether MCQ can truly measure LLM’s capabilities. |
| Approach: | They propose to use multiple choice questions to evaluate large language models (LLMs) to assess their capabilities. |
| Outcome: | The proposed methods show that MCQs are less reliable than LFGQs in terms of expected calibration error. |
Copied to clipboard
| Challenge: | Recent explosion of question-answering datasets and models has increased interest in generalization of models across multiple domains and formats. |
| Approach: | They propose to combine expert agents with a flexible and training-efficient architecture that considers questions, answer predictions, and answer-prediction confidence scores to select the best answer among a list of answer predictions. |
| Outcome: | The proposed model outperforms previous multi-agent and multi-dataset approaches and is highly data-efficient to train and adaptable to any QA format. |
Copied to clipboard
| Challenge: | a primary challenge faced by extractive summarization systems is the lack of annotated data. |
| Approach: | They propose a supervised extractive summarization system that rewards question-answering by identifying salient sequences of words from a document and highlighting them in the text. |
| Outcome: | The proposed system compares with baselines of strong summarization and human assessors on question-answering. |
Copied to clipboard
| Challenge: | Large language models (LLMs) generate information with hallucinations due to uneven retrieval quality and irrelevant contents. |
| Approach: | They propose a decoding strategy which dynamically amplifies knowledge from selected documents during the generation phase. |
| Outcome: | The proposed method outperforms other decoding strategies on ALCE-ASQA, NQ, TQA and PopQA benchmarks. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have improved document understanding performance but generate thousands of visual tokens for a single document image, leading to excessive GPU memory and slower inference times. |
| Approach: | They propose a high-resolution document compression module to generate 324 tokens for a single document image. |
| Outcome: | The proposed module reduces first token latency by more than 50% and improves document comprehension performance. |
Copied to clipboard
| Challenge: | Existing research on streaming video understanding focuses on isolated aspects of visual understanding, but ignores practical deployability under realistic resource constraints. |
| Approach: | They propose a framework to evaluate streaming video understanding capabilities under realistic constraints. |
| Outcome: | StreamingEval benchmarks offline and online video models under a standardized protocol . it evaluates visual encoding efficiency, text decoding latency and task performance . |
Copied to clipboard
| Challenge: | Open-domain complex question-answering systems face challenges in retrieving and reasoning over information that addresses multifaceted queries. |
| Approach: | They propose a method that leverages large language models to guide a Neighborhood Aware Retrieval process. |
| Outcome: | The proposed approach outperforms retrieve-and-reason baselines on two complex QA datasets. |
Copied to clipboard
| Challenge: | Automatic question generation is an increasingly important task that can be applied in educational settings, data augmentation for question-answering (QA), and conversational systems. |
| Approach: | They adapt and apply QAG approaches to generate question-answer pairs given context and look into strategies for error filtering and their effects. |
| Outcome: | The proposed methods can generate question-answer pairs in Portuguese, a widely spoken language that is underrepresented in natural language processing research. |
Copied to clipboard
| Challenge: | Pre-trained language models (PLMs) are often easily deceived by tricky questions such as “How many eyes does the sun have?” . |
| Approach: | They annotate a FalseQA dataset containing 2365 human-written FPQs and find that PLMs are capable of discriminating FPqs by fine-tuning on moderate numbers. |
| Outcome: | The proposed model can discriminate on FPQs by fine-tuning on moderate numbers of examples and generate reasonable explanations for false premise questions. |
Copied to clipboard
| Challenge: | Existing studies on multitask and multilingual learning have shown that learning cross-lingual embeddings can benefit multiple tasks and languages. |
| Approach: | They propose a meta-learning approach to learn interactions between tasks and languages . they also investigate the role of different sampling strategies used during meta-learned model . |
| Outcome: | The proposed model improves on five different tasks and six different languages from the XTREME multilingual benchmark dataset. |
Copied to clipboard
| Challenge: | Existing methods focus on single-step reasoning, ignoring logical dependencies between steps. |
| Approach: | They propose a method that maximizes a structure-based return to facilitate structured reasoning and explanation. |
| Outcome: | The proposed method outperforms state-of-the-art methods on EntailmentBank and STREET benchmarks. |
Copied to clipboard
| Challenge: | Evaluating role-playing capabilities in large language models is challenging due to complex dynamics involved in role-playering. |
| Approach: | They propose a simulation sandbox that generates situational fine-grained character behavior trajectories to enhance LLM performance. |
| Outcome: | The proposed model generates situational fine-grained character behavior trajectories to enhance performance. |
Copied to clipboard
| Challenge: | Language Models excel in understanding textual descriptions of proteins, but struggle to process texts. |
| Approach: | They propose a framework for Protein-to-Text Generation for Text-based Protein Understanding that integrates a PLM as its protein understanding module. |
| Outcome: | The proposed framework surpasses existing baselines and is highly efficient in protein-to-text generation. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown impressive capabilities in many tasks, including natural language understanding and generation. |
| Approach: | They propose a framework for adaptation with self-evaluation to improve selective prediction performance of large language models. |
| Outcome: | The proposed framework outperforms state-of-the-art selective prediction methods on QA datasets and improves the AUACC from 91.23% to 92.63% and AUROC from 74.61% to 80.25%. |
Copied to clipboard
| Challenge: | Existing methods to classify intents are labor-intensive and time-consuming as intents will be diverse and new intents may be involved. |
| Approach: | They propose a zero-shot intent detection problem which aims to detect emerging user intents where no labeled utterances are currently available. |
| Outcome: | The proposed model can discriminate emerging intents when no labeled utterances are available in training data. |
Copied to clipboard
| Challenge: | Existing models that attribute mental states to oneself and others perform poorly on false belief tasks where beliefs differ from reality. |
| Approach: | They propose a temporally informed approach for improving the theory of mind capability of memory-augmented neural models by integrating priors about entities’ minds and tracking their mental states over time through an extended passage. |
| Outcome: | The proposed model improves performance on false belief tasks where beliefs differ from reality, especially when the dataset contains distracting sentences. |
Copied to clipboard
| Challenge: | Language models suffer from poor interpretability and transparency, as well as the intrinsic risk of hallucination and misinformation. |
| Approach: | They propose a statistical framework that assesses how well a query can be answered by an RAG system by capturing the relevance of knowledge. |
| Outcome: | The proposed framework assesses how well a query can be answered by an RAG system by capturing the relevance of knowledge. |
Copied to clipboard
| Challenge: | Existing frameworks that generate single-step reasoning do not improve QA reasoning . |
| Approach: | They propose a framework that strategically constructs and refines sub-questions and their answers (sub-QAs) they argue that sub-QA does not always enhance QA reasoning . |
| Outcome: | The proposed framework can be integrated with existing QA models and benchmarks. |
Copied to clipboard
| Challenge: | Existing automated forecasting studies rely on structured data to predict future events. |
| Approach: | They propose a question-answering task that limits access to unstructured text data . they use a crowdsourced dataset to form a restricted-domain, multiple-choice, question-announcement task . |
| Outcome: | The proposed model achieves 61.0% accuracy on the dataset, which still lags behind human performance by about 19%. |
Copied to clipboard
| Challenge: | Detecting contradictions in texts is often regarded as determining relation between hypothesis and piece of premise. |
| Approach: | They propose a human-annotated dataset to study self-contradictions in long documents . they analyze the capabilities of four open-source and commercially available LLMs . |
| Outcome: | The proposed dataset outperforms open-source LLMs on document-level tasks but struggles with self-contradictions that require more nuance and context. |
Copied to clipboard
| Challenge: | Existing studies have focused on the spatial reasoning capabilities of modern language models (LMs) however, there has been limited research into the spatial thinking capabilities of LMs. |
| Approach: | They propose a question-answering (QA) benchmark for spatial reasoning on natural language text which contains more realistic spatial phenomena not covered by prior work. |
| Outcome: | The proposed method significantly improves LMs' ability on spatial understanding, which in turn helps solve two external datasets, bAbI, and boolQ. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have reshaped the field of natural language processing (NLP) however, fundamental NLP tasks that involve linguistic analysis still play essential roles in the field. |
| Approach: | They propose to use constituency parsing to improve performance of LLMs on deep syntactic parse trees to prompt LLM chunking, filter out low-quality chunks and add remaining chunks to prompts to instruct LLM for parser. |
| Outcome: | The proposed approach improves LLMs' performance on constituency parsing on English and Chinese benchmark datasets. |
Copied to clipboard
| Challenge: | Recent literature reveals that Large Language Models (LLMs) hallucinate intermittently, which impedes their reliability for further utilization. |
| Approach: | They propose a self-detection method to detect which questions an LLM does not know by combining the two components to identify whether the model generates a non-factual response to the question. |
| Outcome: | The proposed method can detect which questions an LLM does not know across factoid question-answering, arithmetic reasoning, and commonsense reasoning tasks. |
Copied to clipboard
| Challenge: | Towards evaluating and improving AI systems in this domain, we propose a mathematical reasoning benchmark based on 23 diversetasks . |
| Approach: | They propose a mathematical reasoning benchmark that includes 23 diverse tasks . they extend the benchmark by collecting task instructions and solutions in the form of Python programs . |
| Outcome: | The proposed model improves on multi-tasking while the best performing model only achieves 60.40%. |
Copied to clipboard
| Challenge: | Pre-trained parsers perform poorly on domain-specific questions, a paper argues . retraining parser with domain- specific questions is expensive, as these require linguistic expertise. |
| Approach: | They propose an automatic labeled domain question generation framework leveraging domain knowledge and seed domain questions. |
| Outcome: | The proposed framework improves state-of-the-art parsers on domain questions. |
Copied to clipboard
| Challenge: | Large language models face vulnerabilities related to the extraction of sensitive information. |
| Approach: | They propose a method to exploit the model's lower-ranked output tokens to extract private information from retrieved documents or training knowledge. |
| Outcome: | The proposed method is effective in both the agentic application privacy extraction setting and the direct training data extraction. |
Copied to clipboard
| Challenge: | Existing Reward Models (RMs) struggle in Retrieval Augmented Generation settings. |
| Approach: | They propose a method that repurposes question-answering datasets into preference pairs that prioritise groundedness over stylistic features. |
| Outcome: | The proposed method surpasses existing RMs trained on larger general corpora with an absolute improvement of +15.5%. |
Copied to clipboard
| Challenge: | Existing rule-based question generation models rely on one or two sentences as input, while long text has posed challenges for sequence to sequence neural models. |
| Approach: | They propose a maxout pointer mechanism with gated self-attention encoder to address the challenges of processing long text inputs for question generation. |
| Outcome: | The proposed model outperforms existing models with sentence-level or paragraph-level inputs pushing the state-of-the-art result from 13.9 to 16.3 (BLEU_4). |
Copied to clipboard
| Challenge: | Existing metrics for evaluating the quality of automatically generated questions are expensive and penalise valid questions that may not have high lexical or semantic similarity to the reference questions. |
| Approach: | They propose a question-answering and span scorer metric based on the answerability of the candidate question given the context. |
| Outcome: | The proposed metric has higher correlation with human judgment without relying on the reference question. |
Copied to clipboard
| Challenge: | Answering questions within business and finance requires reasoning, precision, and a wide-breadth of technical knowledge. |
| Approach: | They propose a benchmark for evaluating models’ ability to reason about realistic financial problems by focusing on question-answering over financial data via program synthesis. |
| Outcome: | The proposed benchmark evaluates models' financial background knowledge, ability to parse financial documents, and capacity to solve complex problems with code. |
Copied to clipboard
| Challenge: | Recent studies have attempted to enhance the performance of large language models (LLMs) in complex question-answering (QA) tasks by combining step-wise planning with external retrieval. |
| Approach: | They propose a framework for enhancing LLMs’ planning capabilities by using planning data derived from knowledge graphs (KGs). |
| Outcome: | The proposed framework improves LLMs’ planning capabilities by using knowledge graphs (KGs) the proposed framework is compared with existing frameworks on multiple datasets and shows that it is effective for large language models. |
Copied to clipboard
| Challenge: | Recent developments in balancing usefulness and safety of large language models raise a critical question . current attacks, especially adversarial ones that manipulate malicious prompts, often aim to manipulate the input . |
| Approach: | They show that LLMs can effectively summarize malicious long documents but often refuse to translate them. |
| Outcome: | The findings highlight a vulnerability in LLMs that can't translate or summarize documents . the study focuses on LLM models, Gemini and GPT-4, which can' be exploited . |
Copied to clipboard
| Challenge: | Existing research on the Somali language information retrieval relies on query translation . lack of digital resources is key obstacle to advancing language technologies . |
| Approach: | They develop an annotated corpus for Somali information retrieval using query expansion technique. |
| Outcome: | The proposed corpus comprises 2335 documents collected from well-known online sites . it can be used for text classification-related tasks and question-answering research purposes. |
Copied to clipboard
| Challenge: | Continued pre-training on paraphrased data has shown empirical promise for enhancing knowledge acquisition, but this approach is costly and unreliable as it relies on external models or manual effort for rewriting. |
| Approach: | They propose formatting-based data augmentation which diversifies documents conveying the same knowledge by altering document formats rather than their content. |
| Outcome: | The proposed methods improve generalization to diverse paraphrased contexts and enhance pre-training and instruction tuning. |
Copied to clipboard
| Challenge: | Existing approaches to complex question-answering (CQA) exhibit uneven performance when questions have different types, harboring inherently different characteristics, e.g., difficulty level. |
| Approach: | They propose a meta-reinforcement learning approach to program induction in CQA to tackle the potential distributional bias in questions. |
| Outcome: | The proposed method achieves state-of-the-art performance on the CQA dataset while using only five trial trajectories for the top-5 retrieved questions in each support set. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly being integrated with social applications . large data sets are limited in their representation of information and do not capture knowledge from the Web . |
| Approach: | They propose a gamified framework that uses collective sensemaking to collect artifacts from 19 different Indian geographic subcultures and benchmark four popular LLMs. |
| Outcome: | The proposed framework is based on 260 participants from 19 different Indian geographic subcultures and shows that it can be used across regional sub-cultures. |
Copied to clipboard
| Challenge: | Existing benchmarks for pre-trained language models are limited to only a few languages . a limited number of tasks are evaluated on non-standardized protocols . |
| Approach: | They propose to aggregate diverse downstream tasks into a benchmark to assess PLMs' qualities . they evaluate 8 pre-trained masked language models on general and biomedical-specific data . |
| Outcome: | The proposed benchmark assesses pre-trained language models on 20 diversified tasks. |
Copied to clipboard
| Challenge: | Multi-agent collaboration among models has shown promise in reasoning tasks but is underexplored in long-form generation tasks like summarization and question-answering. |
| Approach: | They propose a multi-agent multi-model reasoning recipe to improve faithfulness through refinement. |
| Outcome: | The proposed method improves faithfulness and error detection on three summarization datasets and on long-form question-answering tasks. |
Copied to clipboard
| Challenge: | Existing systems for character representation have simplified the problem of representing complex characters via graphs and brief character descriptions. |
| Approach: | They propose a ‘character sheet’ based representation that organizes and filters textual information about characters. |
| Outcome: | The proposed representation organizes and filters textual information about characters and is better and more flexible than previous models. |
Copied to clipboard
| Challenge: | Knowledge graphs are used for information extraction, search engines, question answering, and recommendation systems. |
| Approach: | They propose a French knowledge graph dataset based on RezoJDM. |
| Outcome: | The proposed dataset can be used in many downstream tasks for the French language . it shows that it embeds knowledge graph baselines for link prediction tasks . |
Copied to clipboard
| Challenge: | Existing studies on retrievers and LLMs treat them as separate components . a novel bridge model is proposed to optimize the relationship between the retriever and the LLM . |
| Approach: | They propose a framework that chains together supervised and reinforcement learning to train a bridge model that optimizes the connection between the retriever and the LLM. |
| Outcome: | Empirical results show that the proposed model optimizes the connection between the retriever and the LLM. |
Copied to clipboard
| Challenge: | Open-source vision-language models excel on simple question-answering tasks, but struggle with complex questions that require both perception and reasoning. |
| Approach: | They propose a family of vision-language models that have LeArned to Think wiTh vision spEcialists by offloading perception to state-of-the-art vision models. |
| Outcome: | The proposed model achieves 4-5% gains over baselines across 6 benchmarks covering both perception and reasoning abilities. |
Copied to clipboard
| Challenge: | Retrieval-augmented generation (RAG) enhances the question answering abilities of large language models (LLMs) however, adapting general-purpose RAG systems to specialized fields poses unique challenges due to distribution shifts and limited access to domain-specific data. |
| Approach: | They propose a method that equips large language models with joint capabilities of question answering and question generation for domain adaptation. |
| Outcome: | Experiments on 11 datasets across three different domains verify the efficacy of SimRAG over baselines by 1.2%–8.6%. |
Copied to clipboard
| Challenge: | Currently, the most commonly used confidence score is the likelihood of the generated sequence . different tokens should be weighted differently depending on the context. |
| Approach: | They propose to assign different weights to various tokens using attention values elicited from the base LLM. |
| Outcome: | The proposed model improves the confidence of the predicted sequence probability by assigning weights to tokens based on attention values elicited from the base model. |
Copied to clipboard
| Challenge: | Existing contrastive methods that ignore the context of a large language model (LLM) fail to handle instances that vary in their amount of conflict, with static methods over-adjusting when conflict is absent. |
| Approach: | They propose a fine-grained, instance-level approach called AdaCAD which dynamically adjusts the degree of conflict based on the degree. |
| Outcome: | The proposed approach outperforms baselines and improves factuality of summaries by 6.19. |
Copied to clipboard
| Challenge: | Xu and Peng, 2025) . . SPUR is a comprehensive benchmark for scientific experimental image perception, understanding, and reasoning, comprising 4,264 question-answering (QA) pairs derived from 1,084 expert-curated images. |
| Approach: | They propose to use 4,264 question-answering (QA) pairs derived from 1,084 expert-curated images to evaluate the visual perception of multimodal large language models (MLLMs) . they also propose to utilize cross-panel relation understanding to evaluate MLLM’s ability to decipher intricate cross-panel relations. |
| Outcome: | The proposed model is based on 4,264 question-answering pairs derived from 1,084 expert-curated images. |
Copied to clipboard
| Challenge: | Existing prompt-based debiasing methods exhibit instability due to sensitivity to prompt changes . fine-tuning-based techniques incur substantial computational overhead and catastrophic forgetting . |
| Approach: | They propose a debiasing framework that encodes fairness-related features into separable directions in the hidden activation space. |
| Outcome: | The proposed framework performs inference-time debiasing without requiring retraining or prompt design . it detects bias signatures in activations and then computes debiased steering vectors . the proposed framework is available to download in the u.s. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have transformed natural language processing, but their safety mechanisms remain under-explored in low-resource, multilingual settings. |
| Approach: | They propose a red-teaming approach to probe LLM vulnerabilities in Singapore's diverse linguistic context using a dataset and evaluation framework. |
| Outcome: | The proposed framework systematically probes LLM vulnerabilities in three real-world scenarios including Singlish, Chinese, Malay, and Tamil. |
Copied to clipboard
| Challenge: | Charts provide visual representations of data and are used for analyzing information, addressing queries, and conveying insights to others. |
| Approach: | They propose a chart-specific vision-language Instruction-following dataset with 191K instructions and a pipeline model that extracts chart data tables and inputs them into a LLM. |
| Outcome: | The proposed model can solve a wide range of chart-related tasks, achieving state-of-the-art results on four tasks. |
Copied to clipboard
| Challenge: | Eye movement features are considered to be direct signals reflecting human attention distribution with a low cost to obtain, inspiring researchers to augment language models with eye-tracking (ET) data. |
| Approach: | They select first fixation duration (FFD) and total reading time (TRT) as the cognitive signals to guide Transformer attention in question-answering tasks. |
| Outcome: | The proposed models improve but compromise stability when augmenting with ET data. |
Copied to clipboard
| Challenge: | Dual encoders have been used for question-answering and information retrieval tasks with good results. |
| Approach: | They propose to use two different versions of dual encoders for QA retrieval tasks . they propose to share parameters in projection layers between two encoder towers . |
| Outcome: | The proposed architectures outperform SDE and ADE on QA retrieval tasks. |
Copied to clipboard
| Challenge: | Using simulated feedback, our system (called TeachMe) continually improves with time, and without model retraining. |
| Approach: | They propose to augment a QA model with a dynamic memory of user feedback, containing user-supplied corrections toerroneous model beliefs that users identify during interaction. |
| Outcome: | The proposed system improves with time and without model retraining, and with real users, by 15% on a hidden test set after teaching. |
Copied to clipboard
| Challenge: | Large language models have shown significant promise in question-answering tasks . noisy reference documents hinder performance of LLMs, causing disproportionate attention to irrelevant content . |
| Approach: | They propose an adaptive large language model that allocates disproportionate attention to irrelevant documents . they use transformers to train the model and integrate it into pre-trained Transformer blocks . |
| Outcome: | The proposed model outperforms state-of-the-art models on noisy-context benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for open attribute value extraction for emerging entities are noisy or incomplete, even missing. |
| Approach: | They propose a knowledge-guided reinforcement learning framework for open attribute value extraction for emerging entities. |
| Outcome: | The proposed framework outperforms baselines by 16.5 - 27.8%. |
Copied to clipboard
| Challenge: | Existing tabular QA models are lacking in understanding their robustness on scientific information. |
| Approach: | They propose a dataset to assess the robustness of tabular QA models on scientific hybrid tabular data. |
| Outcome: | The proposed model performs well on scientific tables and text, while the best score is 0.462. |
Copied to clipboard
| Challenge: | Using a dataset of tweets and Reddit, we investigate the public opinion on cryptocurrency and bitcoin on Twitter and RedDit. |
| Approach: | They create a dataset to investigate the public opinion on cryptocurrency and bitcoin on Twitter and Reddit. |
| Outcome: | The proposed dataset contains gold standard and silver standard labels and a question-answering sub-corpus. |
Copied to clipboard
| Challenge: | Long-context large language models miss important information in the middle of context documents . a recent study shows that LLMs can be used for document-based QA tasks . |
| Approach: | They propose a prompt-based method called *reprompting* and *in-context retrieval* to alleviate this effect in document-based QA. |
| Outcome: | The proposed method improves QA accuracy on documents up to 80k tokens in length. |
Copied to clipboard
| Challenge: | Multimodal Retrieval Augmented Generation (MMRAG) is a powerful approach to question-answering over multimodal documents. |
| Approach: | They propose a synthetic data generation framework that leverages interplay between a retriever, large language model and large multimodal model to generate question and answer pairs directly from multimodal documents. |
| Outcome: | The proposed framework generates question and answer pairs from 1024 questions over Wikipedia documents and evaluates state-of-the-art models using it. |
Copied to clipboard
| Challenge: | Existing models cannot reason about or explain learned science concepts in novel contexts, despite transformer-based progress in question-answering and scientific text processing . |
| Approach: | They propose a benchmark to test agents’ scientific reasoning abilities in a new interactive text environment at the level of a standard elementary school science curriculum. |
| Outcome: | The proposed model outperforms a model trained for 100k steps in a standard elementary school science curriculum. |
Copied to clipboard
| Challenge: | Existing Paper2Video systems are monolingual and often rely on single-pass pipelines. |
| Approach: | They propose a multilingual agentic Paper2Video system that decomposes the task into planning, audience-oriented critique, layout-aware slide generation, and multilingual figure interpretation. |
| Outcome: | The proposed system improves question-answering accuracy relative to previous systems while maintaining affordable cost and latency. |
Copied to clipboard
| Challenge: | Visual word sense disambiguation (VWSD) is a challenging task involving multiple candidates . context given for an ambiguous word is minimal, most often limited to a single word . |
| Approach: | They propose to use large language models to enhance given phrases and resolve ambiguity related to the target word. |
| Outcome: | The proposed frameworks improve the image representation of ambiguous words among candidates and achieve competitive ranking results. |
Copied to clipboard
| Challenge: | Existing question-answering benchmarks fail to evaluate SLMs’ knowledge understanding due to their inability to support end-to-end speech evaluation and account for varied input audio conditions. |
| Approach: | They propose a new question-answering benchmark that assesses SLMs’ knowledge understanding through pure speech interactions. |
| Outcome: | The proposed benchmark maintains speech format for both inputs and outputs, evaluates model robustness across diverse input audio conditions, and pioneers the assessment of complex tasks like mathematical reasoning in spoken format. |
Copied to clipboard
| Challenge: | Existing methods to generate RAG documents require knowledge of the target RAG system’s internal composition and implementation details, whereas black-box methods are unable to utilize interactive information. |
| Approach: | They propose a RIPRAG attack framework that treats the target RAG system as a black box and leverages a Reinforcement Learning from Black-box Feedback (RLBF) method to optimize the generation model for poisoned documents. |
| Outcome: | The proposed method achieves an attack success rate (ASR) improvement of up to 0.72 compared to baseline methods. |
Copied to clipboard
| Challenge: | Meeting transcripts are a promising domain for natural language tasks . lack of annotated data impedes research on other important tasks in this domain . |
| Approach: | They propose an extractive QA dataset comprising questions asked by meeting participants and corresponding responses. |
| Outcome: | The proposed dataset extracts questions asked by meeting participants and corresponding responses from transcripts. |
Copied to clipboard
| Challenge: | Existing studies have focused on Knowledge Graph Completion as an end in itself, neglecting its potential impact on subsequent applications. |
| Approach: | They propose a benchmark to assess the impact of representative KGC methods on Knowledge Graph Question Answering (KGQA) they use a knowledge graph with 3 million triplets across 5 distinct domains to evaluate their results. |
| Outcome: | The proposed benchmark compares four well-known methods with two state-of-the-art systems to assess the impact of incomplete knowledge graphs on KGQA. |
Copied to clipboard
| Challenge: | Entity disambiguation (ED) is crucial in natural language processing tasks such as question-answering and information extraction. |
| Approach: | They propose a method to reduce computational overhead on overshadowed entities by addressing shortcut learning. |
| Outcome: | The proposed method achieves state-of-the-art performance without compromising inference speed. |
Copied to clipboard
| Challenge: | Large language models have shown a powerful ability for text generation, but undesired behaviors such as toxicity and hallucinations can manifest. |
| Approach: | They propose to formalize text generation as a future-constrained generation problem to minimize undesirable behaviors and enforce faithfulness to instructions. |
| Outcome: | The proposed approach is effective across three tasks, including keyword-constrained generation, toxicity reduction, and factual correctness in question-answering. |
Copied to clipboard
| Challenge: | lack of interpretability is a growing impediment to widespread use of large language models . a new approach to solve this problem is to add a rational layer on top of the LLM . |
| Approach: | They propose to add a rational layer to the large language models to make model beliefs explicit . they also propose to identify and minimize contradictions in the model belief graph . |
| Outcome: | a new approach improves consistency without harming overall answer accuracy . the proposed approach makes model beliefs explicit and resolves inconsistencies . |
Copied to clipboard
| Challenge: | Existing approaches to document understanding are limited due to limited context length or fail to fully leverage multi-modal information. |
| Approach: | They propose a multi-agent framework for long-context document understanding that imitates human reading practice. |
| Outcome: | The proposed framework surpasses human-level benchmarks on long-context document understanding while maintaining a short context length. |
Copied to clipboard
| Challenge: | Question answering (QA) is a fundamental task in the field of Natural Language Processing (NLP). |
| Approach: | They propose a database querying and reasoning dataset for question answering that is designed to accommodate sequential questions and multi-hop queries. |
| Outcome: | The proposed dataset better mirrors the dynamics of real-world information retrieval and analysis with a particular focus on the financial reports of US companies. |
Copied to clipboard
| Challenge: | Existing studies on text-discriminating properties of semi-parametric models have not been done on non-parameter models. |
| Approach: | They propose an inference-phase approach that incorporates a neighborhood search into a model to enhance the capacity of a pre-trained parametric text classifier. |
| Outcome: | The proposed model improves performance on eight SuperGLUE tasks, three adversarial natural language inference datasets, 11 question-answering (QA) datasets and two sentiment classification datasets. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have been successful in understanding language and processing text, but their cost prohibits their practical applications. |
| Approach: | They propose a multi-agent collaboration method that breaks down lengthy documents into smaller, more manageable chunks and organizes the member agents to read their assigned chunks. |
| Outcome: | The proposed method achieves 16.42% and 1.63% accuracy gains over existing models on single-hop and multi-hop QA settings. |
Copied to clipboard
| Challenge: | Existing story reading systems fail to capture the nuances of how education experts think when conducting interactive story reading activities. |
| Approach: | They propose to use existing question-answering (QA) datasets to capture experts' annotations and thinking process to construct a story-based annotation framework. |
| Outcome: | The proposed framework captures experts’ annotations and thinking process and can be used to generate 5, 868 expert-annotated QA pairs with real-world knowledge. |
Copied to clipboard
| Challenge: | Existing tools for extracting information about net zero and emission reduction targets have not been used to assess the vast amounts of information about sustainability commitments made by public and private actors. |
| Approach: | They propose a data set and train and release a natural language classifier to detect whether a text contains a net zero or reduction target. |
| Outcome: | The proposed model can be combined with conventional Q&A models to analyze the ambitions displayed in net zero and reduction targets. |
Copied to clipboard
| Challenge: | Existing solutions for long-form question-answering (LFQA) use chain-of-thought (CoT) with retrieval-augmented generation (RAG). |
| Approach: | They propose to integrate chain-of-thought (CoQ) with retrieval-augmented generation to improve answer comprehensiveness and verifiability. |
| Outcome: | The proposed approach outperforms ChatGPT baselines while maintaining efficiency. |
Copied to clipboard
| Challenge: | Existing methods for hallucination detection rely on self-consistency check alone . prominent LMs exhibit a tendency to produce exceedingly confident, but erroneous, assertions . |
| Approach: | They propose a sampling-based method that expands on the principle of self-consistency checking to detect hallucinations at question-level and model-level. |
| Outcome: | The proposed method outperforms the state of the art in detecting non-factual and factual statements across multiple question-answering and open-domain generation benchmarks. |
Copied to clipboard
| Challenge: | Recent instruction-finetuned large language models (LMs) have shown notable performances in various tasks, such as question-answering. |
| Approach: | They propose to use unlabeled test data to transfer smaller language models with limited knowledge. |
| Outcome: | The proposed strategy shows significant performance improvements on benchmark QA datasets with higher robustness across diverse prompts, enabling LMs to stay stable. |
Copied to clipboard
| Challenge: | Recurrent exchange of model updates in FL can result in prohibitively high communication costs, hindering the distributed learning process. |
| Approach: | They propose a federated fine-tuning framework that uses a round-robin segment sharing scheme to reduce network bandwidth and adaptive sparsification methods tailored to LoRA’s training dynamics. |
| Outcome: | The proposed framework reduces communication overhead without compromising performance on question-answering and value-alignment tasks. |
Copied to clipboard
| Challenge: | MATCHA is an automatic metric that rewards semantic agreement with a reference and penalizes contradictions. |
| Approach: | They introduce a metric that jointly rewards semantic agreement with a reference and penalizes contradictions. |
| Outcome: | The proposed metric outperforms popular metrics on eight public benchmarks compared with human annotations on question-answering, image caption generation, natural language inference, summarization, and semantic textual similarity tasks. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) demonstrate strong visual question answering (VQA) capabilities but are shown to hallucinate. |
| Approach: | They propose three confidence-based methods to enhance LVLMs' perception . they propose probabilistic and consistency-based signals are more reliable indicators . |
| Outcome: | Experiments on three LVLMs across three VQA datasets show that LVLs possess a reasonable perception level but there is room for improvement. |
Copied to clipboard
| Challenge: | Recent advances in large language models have shown promising ability to perform commonsense reasoning. |
| Approach: | They propose a two-dimensional analysis framework that incorporates token back-tracing and token decoding to uncover how LLMs conduct factual knowledge recall. |
| Outcome: | The proposed framework shows that LLMs lack relevant knowledge but struggle to select the most accurate information based on context during the retrieval and rerank phase. |
Copied to clipboard
| Challenge: | Temporal relation annotation in the clinical domain is crucial but challenging due to its workload and the medical expertise required. |
| Approach: | They propose an annotation method that integrates event start-points ordering and question-answering as the annotation format. |
| Outcome: | The proposed method achieves a 0.72 F1 score and enables collaboration among medical experts and non-experts. |
Copied to clipboard
| Challenge: | a recent study has shown that short video understanding is not trivial due to the need for long-range temporal reasoning capabilities. |
| Approach: | They propose a language-based short- and long-range question-answering framework LLoVi . they propose 'multi-round summarization prompt' that asks the LLM to summarize the captions . |
| Outcome: | The proposed framework outperforms the state-of-the-art on the EgoSchema dataset and to grounded VideoQA. |
Copied to clipboard
| Challenge: | Existing work assesses models’ generalization capabilities through the lens of performance on out-of-distribution (OOD) datasets. |
| Approach: | They challenge this assumption by comparing OOD evaluations with failure modes documented in existing question-answering (QA) models. |
| Outcome: | The proposed evaluations show that the models' generalization capabilities are under-performing on out-of-distribution datasets, while others are underperforming on in-difference datasets. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have become increasingly prevalent in the field of Natural Language Processing (NLP), achieving unprecedented performance across linguistic tasks. |
| Approach: | They propose a framework to quantify and analyze context-driven over-refusal . they find that over-fusals depend on the task, system prompts, model family, and the number of retrieved documents. |
| Outcome: | The proposed framework quantifyes and analyzes the concept of context-driven over-refusal on two public corpora. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown remarkable performance on question-answering tasks due to their superior capabilities in natural language understanding and generation. |
| Approach: | They propose a structured taxonomy that categorizes the methodology of synthesizing LLMs and knowledge graphs for QA according to the categories of QA and the KG’s role when integrating with LLM. |
| Outcome: | The proposed taxonomy categorizes the methods according to the categories of QA and the KG’s role when integrating with LLMs. |
Copied to clipboard
| Challenge: | Limited availability of multilingual text corpora for pretraining results in poor performance on downstream tasks due to undertrained representation spaces for languages other than English. |
| Approach: | They propose a method that integrates source and target language representations within low-rank (LoRA) adapters using lightweight linear transformations to enhance representation quality and transfer performance for languages other than English. |
| Outcome: | The proposed method improves representation quality and performance for languages other than English while maintaining parameter efficiency. |
Copied to clipboard
| Challenge: | Large language models (LLMs) perform well on well-posed factual queries, yet standard question-answering (QA) benchmarks remain far from solved. |
| Approach: | They propose an LLM-based classifier to identify underspecified questions and apply it to several widely used QA datasets. |
| Outcome: | The proposed classifier detects underspecified questions in QA datasets and significantly improves on them. |
Copied to clipboard
| Challenge: | Multi-document (MD) processing is crucial for LLMs to handle real-world tasks such as summarization and question-answering across large sets of documents. |
| Approach: | They propose a framework that generates high-quality synthetic MD instruction data over sets of articles via targeted prompts. |
| Outcome: | MDCure generates high-quality synthetic MD instruction data over sets of articles . evaluations show it improves over pre-trained models by up to 75.1% . |
Copied to clipboard
| Challenge: | Large language models (LLMs) are being used for context-grounded tasks like summarizing meetings and answering doctors' questions. |
| Approach: | They propose a technique to provide explanations for context-grounded text generation by assigning scores to parts of the context to quantify their influence on the model output. |
| Outcome: | The proposed framework can provide more faithful explanations of generated output than available alternatives, including LLM self-explanations. |
Copied to clipboard
| Challenge: | Relevance emphasizes the aboutness of a result to a query, while utility refers to the result’s usefulness or value to an information seeker. |
| Approach: | They propose an Iterative utiliTy judgmEnt fraMework to promote each step in Retrieval-Augmented Generation (RAG) they propose to use relevance ranking, utility judgments, and answer generation to prioritize high-utility results over low-utilitity results. |
| Outcome: | The proposed framework improves relevance, ranking, and answer generation on retrieval (TREC DL, WebAP), utility judgment task (GTI-NQ), and factoid question-answering (NQ) datasets. |
Copied to clipboard
| Challenge: | Existing approaches to reinforcement learning from human feedback (RLHF) require expensive human-annotated datasets and proprietary models like GPT-4 to annotate preference pairs. |
| Approach: | They propose a self-synthetic framework for LLM alignment where all training data, including prompts (i.e., user queries), responses, and preferences, are generated by the model itself. |
| Outcome: | The proposed framework enhances the model’s chat capabilities on standard benchmarks like AlpacaEval 2.0 while maintaining strong performance on downstream objective tasks. |
Copied to clipboard
| Challenge: | Existing deep research frameworks lack adequate evaluation procedures and stage-specific protections. |
| Approach: | They propose a framework with open-domain evaluation and a stage-wise safety benchmark to address this oversight. |
| Outcome: | The proposed framework improves defense success rates by 16.53% while reducing over-refusal rates to approximately 6%. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly adopted as scalable judges for open-ended generation, yet how they form judgments remains insufficiently understood. |
| Approach: | They show that exposing reasoning influences LLM-based judgment . they also show that reasoning fluency and factuality critically shape judgment outcomes . |
| Outcome: | Empirical results show that the presence of reasoning significantly alters judgment behavior . stronger judges exhibit more selective behavior and achieve higher judgment accuracy . |