Papers by Caiming Xiong
Copied to clipboard
| Challenge: | Large language models (LLMs) acquire a wide range of abilities and abilities, but their behavior does not align with human preferences. |
| Approach: | They propose to minimize a forward Kullback–Leibler divergence from a target policy to a parameteric policy instead of a reverse KL as in RLHF methods. |
| Outcome: | The proposed method can learn an aligned policy by minimizing a forward Kullback–Leibler divergence from a target policy to a parameteric policy instead of a reverse KL as in RLHF methods. |
Copied to clipboard
| Challenge: | Cross-lingual transfer of language models trained on high-resource languages such as English has been limited due to the high cost of obtaining non-English conversational data. |
| Approach: | They introduce a parallel and large-scale multilingual conversation dataset that is used for cross-lingual alignment pretraining by translating the English-only Schema-Guided Dialogue dataset into 105 other languages. |
| Outcome: | The proposed model performs well on slot-filling and intent classification tasks, and is able to perform well in other languages. |
Copied to clipboard
| Challenge: | Existing studies on text summarization factual consistency are divided into two categories . entailment-based and question answering-based metrics are the most efficient . |
| Approach: | They propose an optimized QA-based metric that improves factual consistency by 14% . they compare entailment-based and QA metrics to find the best fit . |
| Outcome: | The proposed metric outperforms the best performing entailment-based metric on the SummaC factual consistency benchmark. |
Copied to clipboard
| Challenge: | Using a model to generate summary sketches, we improve abstractive dialogue summarization quality and enable granularity control. |
| Approach: | They propose a model that generates a preliminary summary sketch and a strategy to control granularity. |
| Outcome: | The proposed model achieves state-of-the-art on the largest dialogue summarization corpus with as high as 50.79 in ROUGE-L score. |
Copied to clipboard
| Challenge: | Recent work shows that neural QA models are sensitive to adversarial inputs. |
| Approach: | They propose a sentence selector to select the minimal set of sentences to feed into a QA model. |
| Outcome: | The proposed system reduces training time and inference time by up to 13 times . it is comparable to or better than the state-of-the-art on SQuAD, NewsQA, TriviaQA and SQu AD-Open . |
Copied to clipboard
| Challenge: | Existing judge models are largely trained with supervised finetuning on small data scales to perform limited types of evaluation tasks, limiting generalization. |
| Approach: | They propose to train judge models at large data scales with direct preference optimization . they use four training tasks to form three types of preference pairs targeting different aspects of evaluation . |
| Outcome: | The proposed model outperforms GPT-4o and other similar models on 13 benchmarks. |
Copied to clipboard
| Challenge: | Existing neural question generation approaches focus on short factoid type of answers. |
| Approach: | They propose a neural question generator that trains a single generative model by combining multiple question types with different answer types. |
| Outcome: | The proposed model outperforms existing models in both seen and unseen domains and can generate questions with different cognitive levels when conditioned on different answer types. |
Copied to clipboard
| Challenge: | BRIDGE is a powerful sequential architecture for cross-modal semantic parsing . BRidege captures cross-modal dependencies between natural language questions and relational databases . |
| Approach: | They propose a sequential architecture that captures cross-modal dependencies between questions and relational databases in cross-DB semantic parsing. |
| Outcome: | The proposed architecture performs well on the well-studied Spider benchmark (65.5% dev, 59.2% test). |
Copied to clipboard
| Challenge: | Existing methods for representation learning of text are masked language modeling (MLM) a language model is trained to learn universal contextual embeddings, which are fine-tuned on a down-stream task. |
| Approach: | They propose a self-critic pretraining transformer for representation learning of text . they demonstrate improved sample-efficiency and improved performance over strong baselines . |
| Outcome: | The proposed model improves sample-efficiency and performance over strong baselines. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly being used for reasoning intensive tasks. |
| Approach: | They propose an algorithm that trains judges to be robust to positional biases . they also propose a benchmark that evaluates judges in diverse reasoning settings . |
| Outcome: | The proposed algorithm outperforms GPT-4o and the next best small judge by 6.7% and 9% on ReasoningJudgeBench and JudgeBench. |
Copied to clipboard
| Challenge: | Existing benchmarks for logical reasoning in large language models lack language naturalness or limited complexity. |
| Approach: | They propose to use first-order logic annotations to evaluate logical reasoning capabilities of large language models. |
| Outcome: | The proposed dataset evaluates the FOL reasoning ability of supervised fine-tuning on medium-sized language models. |
Copied to clipboard
| Challenge: | a growing number of academic articles are shared daily, making it difficult to keep up with the latest findings. |
| Approach: | They propose a task of disentangled paper summarization which generates separate summaries for papers and contexts to make it easier to identify key findings shared in articles. |
| Outcome: | The proposed task is more useful than traditional scientific paper summarization. |
Copied to clipboard
| Challenge: | Vision Language Models struggle with visual arithmetic, seemingly simple tasks like object counting or length comparison, which are essential for relevant complex tasks like chart understanding and geometric reasoning. |
| Approach: | They propose a novel post-training strategy inspired by Piaget’s theory of cognitive development that trains VLMs to recognize invariant properties under visual transformations. |
| Outcome: | The proposed approach outperforms supervised fine-tuning methods while requiring 60% less training data. |
Copied to clipboard
| Challenge: | Prior work on document-level simplification has focused on sentence-level edits, while many desirable edits require document- level context. |
| Approach: | They propose a dataset that reconstructs the document-level editing process from English Wikipedia to paired Simple Wikipedia articles. |
| Outcome: | The proposed dataset reconstructs the document-level editing process from English Wikipedia (EW) articles to paired Simple Wikipedia (SEW) pages. |
Copied to clipboard
| Challenge: | Existing infrastructure for efficient agentic data processing and model training remains underdeveloped. |
| Approach: | They propose a lightweight and extensible data and training framework for large action models . they propose to unify diverse agent trajectories using Unified Format 2.0 . |
| Outcome: | The proposed framework shows 9 higher throughput than existing frameworks and performs well across public and realistic agent benchmarks. |
Copied to clipboard
| Challenge: | Existing fact-checking models trained on non-dialogue data fail to perform well on this task. |
| Approach: | They propose a task of fact-checking in dialogue to improve fact- checking performance . they propose to use an annotated conversational claim and Wikipedia snippets as evidence . |
| Outcome: | The proposed task improves fact-checking performance in dialogue. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have remarkable capabilities, but unreliability remains a barrier to deployment in high-stakes domains. |
| Approach: | They propose to transform uncertainty from a passive diagnostic metric to an active control signal guiding real-time model behavior. |
| Outcome: | The proposed model evolution from passive diagnostic metric to active control signal is critical for high-stakes applications. |
Copied to clipboard
| Challenge: | Cananical Correlation Analysis (CCA) of the internal representations of a pre- trained, multilingual BERT model reveals that the model partitions representations for each language rather than using a common, shared, interlingual space. |
| Approach: | They propose to use a multilingual BERT model to partition representations for each language rather than using a common, shared, interlingual space. |
| Outcome: | The results show that the model partitions representations for each language rather than using a common, shared, interlingual space. |
Copied to clipboard
| Challenge: | Existing studies focus on sentence-level inference, which limits its application in downstream NLP problems. |
| Approach: | They propose to construct a large-scale dataset for document-level NLI that can be used to study NLP problems. |
| Outcome: | The proposed model performs well on popular sentence-level benchmarks and generalizes well to out-of-domain NLP tasks that rely on inference at document granularity. |
Copied to clipboard
| Challenge: | Recent advances in efficient attention mechanisms have led to the expansion of the context length of large language models. |
| Approach: | They propose a procedure to synthesize Haystacks of documents and generate a summary that identifies relevant insights and precisely cites the source documents. |
| Outcome: | The proposed evaluation can score summaries on Coverage and Citation . the proposed evaluation lags human performance estimates by 10+ points on SummHay . |
Copied to clipboard
| Challenge: | Existing research lacks direct access to such data, making benchmarking difficult due to privacy concerns. |
| Approach: | They propose a synthetic data pipeline that generates realistic user profiles and private documents and a benchmark to evaluate models' ability to understand personal information. |
| Outcome: | The proposed pipeline generates realistic user profiles and private documents, enabling PersonaBench, a benchmark for evaluating models’ ability to understand personal information. |
Copied to clipboard
| Challenge: | Prior work focused on attention mechanisms to model complex interactions in visual dialog . a new framework for visual dialog is based on pretrained BERT language models . |
| Approach: | They propose a framework for a vision-dialog Transformer that leverages pretrained BERT language models for Visual Dialog tasks. |
| Outcome: | The proposed framework achieves the top position on the visual dialog leaderboard without pretraining on external vision-language data. |
Copied to clipboard
| Challenge: | Existing methods to detect spoken language deteriorate drastically in discriminating the few-shot intents. |
| Approach: | They propose a composed variational natural language generator (CLANG) that generates natural examples for few-shot intents in imbalanced scenarios. |
| Outcome: | The proposed model achieves state-of-the-art on two real-world intent detection datasets. |
Copied to clipboard
| Challenge: | a data augmentation technique is used to augment data, but it has two drawbacks. |
| Approach: | They propose a new mixup paradigm that generates new points scattered throughout the whole mini-batch. |
| Outcome: | The proposed model improves the performance of NLP tasks while using different ratios of training data. |
Copied to clipboard
| Challenge: | a new model for parallel sentence retrieval can be used to align parallel sentences in multilingual corpora . a faithful aligner can help narrow down the candidate pool without having to deal with an enormous search space . |
| Approach: | They propose a model that can be trained on only one language pair and transfers to low-resource languages with negligible degradation in performance. |
| Outcome: | The proposed model outperforms the previous model on the Tateoba dataset by 8.0 points in accuracy and using less than 0.6% of their parallel data. |
Copied to clipboard
| Challenge: | Existing methods on understanding the capabilities of LLMs in logical reasoning rely on binary entailment classification or synthetically derived rationales. |
| Approach: | They propose to annotate a human-annotated dataset consisting of diverse and complex reasoning chains for a set of realistic logical reasoning stories also written by humans. |
| Outcome: | The proposed model outperforms existing methods on understanding the capabilities of LLMs in logical reasoning by 10% or more. |
Copied to clipboard
| Challenge: | Xu et al., 2017): a dataset for cross-domain semantic parsing in context with 4,298 question sequences. |
| Approach: | They present a dataset for cross-domainSemanticParsing inContext that consists of 4,298 coherent question sequences. |
| Outcome: | The proposed dataset demonstrates that it has greater semantic diversity and can be generalized to unseen domains due to its cross-domain nature and the unseened databases at test time. |
Copied to clipboard
| Challenge: | Abstractive text summarization models do not capture the abstractive nature of high quality summaries. |
| Approach: | They propose to decompose a decoder into a contextual network and a pretrained language model that incorporates prior knowledge about language generation. |
| Outcome: | The proposed model achieves comparable results to state-of-the-art models, based on ROUGE scores and human evaluations, while producing a significantly higher level of abstraction. |
Copied to clipboard
| Challenge: | coding tasks require generated code to be fully executable and functionally correct . current agentic approaches struggle with multi-stage planning, generating, and debugging . |
| Approach: | They propose a framework for LLM agents to efficiently explore the search space in different stages of the code generation process. |
| Outcome: | The proposed framework achieves top results on 7 code generation benchmarks and a 31.9% solving rate on the SWEBench benchmark. |
Copied to clipboard
| Challenge: | Data-to-text annotations can be costly when dealing with tables with nontrivial structures. |
| Approach: | They propose a procedure for extracting semantic triples from tables that encodes their structures by exploiting table headers and table title. |
| Outcome: | The proposed method exploits the semantic dependencies between table headers and title to extract semantic triples from tables. |
Copied to clipboard
| Challenge: | Modern natural language generation paradigms require a decoding strategy to obtain quality sequences out of the model. |
| Approach: | They propose a deterministic search algorithm balancing quality and diversity . they investigate the vanilla best-first search algorithm and propose k-k search algorithm. |
| Outcome: | The proposed algorithm is parameter-free, lightweight, efficient, and easy-to-use. |
Copied to clipboard
| Challenge: | Existing studies on DA tagging focus on human-human social conversations, which is less applicable for task-oriented setting. |
| Approach: | They propose a controllable mechanism that augments text input by leveraging the pre-trained Mask token from BERT model. |
| Outcome: | The proposed mechanism augments text input by leveraging the pre-trained Mask token from BERT model. |
Copied to clipboard
| Challenge: | Existing generative question answering models that leverage passage retrieval with a pre-trained transformer are not effective for multihop QA. |
| Approach: | They propose a generative approach that explicitly models the reasoning process to resolve the answer for multi-hop questions by encoding cross-passage interactions. |
| Outcome: | The proposed model improves on two multi-hop QA datasets and is interpretable. |
Copied to clipboard
| Challenge: | Recent work has demonstrated that deployed NLP models can be stolen by adversaries by querying victim models with gibberish input data that consists of random sequences of words. |
| Approach: | They propose to extract a local copy of a monolingual victim model from an API and query it with gibberish input data paired with the victim's labels. |
| Outcome: | The extracted model learns the task from the monolingual victim, but it generalizes far better than the victim to several other languages. |
Copied to clipboard
| Challenge: | Existing methods to learn correspondence between visual segments and texts require temporal coordinates for training, which leads to high costs of annotation. |
| Approach: | They propose weakly supervised language localization networks to detect events in untrimmed videos . they train with only video-sentence pairs without accessing to temporal locations of events . |
| Outcome: | Experiments on ActivityNet Captions and DiDeMo show that WSLLN performs state-of-the-art. |
Copied to clipboard
| Challenge: | DialogStudio is the largest and most diverse collection of dialogue datasets . existing datasets lack diversity and comprehensiveness, authors say . |
| Approach: | They introduce DialogStudio: the largest and most diverse collection of dialogue datasets . DialogStuio aggregates more than 80 diverse dialogue dataset . |
| Outcome: | a new dataset is created to improve the quality and diversity of dialogue datasets . DialogStudio is the largest and most diverse collection of dialogue data . |
Copied to clipboard
| Challenge: | Question generation models are often evaluated with standardized NLG metrics that are based on n-gram overlap. |
| Approach: | They propose to use QGen to help teachers automate the generation of reading comprehension quizzes by comparing n-gram overlap with BLEU to compare system-generated questions with heldout human-written references. |
| Outcome: | The best model had only 68.4% of its questions accepted by the ten teachers who participated in the study. |
Copied to clipboard
| Challenge: | Compared to neural systems, automatic metrics should be interpretable and provide intuitive insights into system performance and output quality. |
| Approach: | They propose to use a two-stage evaluation pipeline to extract basic information units from one text sequence and check the extracted units in another sequence. |
| Outcome: | The proposed metrics can provide high interpretability at both the fine-grained unit level and summary level, and one-stage metrics that achieve a balance between efficiency and interpretability. |
Copied to clipboard
| Challenge: | Generating SQL queries from user utterances is an important task to help end users acquire information from databases. |
| Approach: | They propose a context-dependent text-to-SQL generation task that edits previous queries . they use an utterance-table encoder and a table-aware decoder to incorporate context . |
| Outcome: | The proposed model is flexible to change individual tokens and robust to error propagation. |
Copied to clipboard
| Challenge: | Dense neural text retrieval has achieved promising results on open-domain Question Answering (QA) current dense retrievers require splitting documents into short passages that usually contain local, partial and sometimes biased context, and may yield inaccurate and misleading hidden representations, thus deteriorating the final retrieval result. |
| Approach: | They propose a hierarchical framework which can generate accurate dense representations of passages by utilizing both macroscopic semantics in the document and microscopic specific to each passage. |
| Outcome: | The proposed framework significantly outperforms the original dense passage retriever and helps an end-to-end QA system outperfect the strong baselines on multiple open-domain QA benchmarks. |
Copied to clipboard
| Challenge: | Existing benchmarks fail to assess large language models’ format-following proficiency adequately. |
| Approach: | They propose a benchmark to evaluate large language models' ability to follow complex, domain-specific formats. |
| Outcome: | The proposed framework evaluates large language models' ability to follow complex, domain-specific formats across open-source and closed-source models. |
Copied to clipboard
| Challenge: | Existing approaches to answer user questions are limited in their decision making due to struggles in extracting question-related rules and reasoning about them. |
| Approach: | They propose a conversational machine reading framework that uses a Explicit Memory Tracker to track whether conditions in the rule text have already been satisfied to make a decision. |
| Outcome: | The proposed framework achieves state-of-the-art on the ShARC benchmark and is more interpretable by visualizing the entailment-oriented reasoning process as the conversation flows. |
Copied to clipboard
| Challenge: | Prompt tuning is an efficient method for adapting large language models, but it is difficult and expensive to identify the source task that provides optimal prompts. |
| Approach: | They propose to learn a shared latent space which captures a set of basis skills from a mixture of source tasks and then transfer them to target tasks. |
| Outcome: | The proposed method outperforms previous methods on NLI, sentence completion, QA, conference resolution, word sense disambiguation and on various model scales. |
Copied to clipboard
| Challenge: | Existing calibration methods rescale posterior distributions of classifiers after training. |
| Approach: | They propose to use a noise contrastive estimation technique to train an energy-based model during finetuning of pretrained text encoders. |
| Outcome: | The proposed model can reach a better calibration competitive to strong baselines with little or no loss in accuracy. |
Copied to clipboard
| Challenge: | Existing KBQA approaches struggle with generalization of unseen KB schema items . Rank-and-generate approach solves coverage issue with strong generalization . |
| Approach: | They propose a Rank-and-Generate approach for KBQA that uses a generation model to generalize to unseen KB schema items. |
| Outcome: | The proposed approach outperforms the prior state-of-the-art on GrailQA and WebQSP datasets. |
Copied to clipboard
| Challenge: | Recent advances in off-policy reinforcement learning methods that use offline data as against a simulator have proven to be sample efficient. |
| Approach: | They propose a batch-RL framework for ToD policy learning: Causal-aware Safe Policy Improvement (CASPI) that uses a mechanism to learn fine-grained reward that captures intention behind human response and offers guarantee on dialogue policy’s performance against a baseline. |
| Outcome: | The proposed framework outperforms the current state of the art on an end-to-end dialogue task using a multiwoz2.0 dataset. |
Copied to clipboard
| Challenge: | Existing table-based question answering datasets lack advanced information-based questions that require reasoning and integration of information pieces retrieved from structured knowledge sources. |
| Approach: | They propose a dataset with 10K Wikipedia-based table, question, free-form answer, supporting table cells pairs that can be used to generate an answer. |
| Outcome: | The proposed dataset has 10K Wikipedia-based table, question, free-form answer, supporting table cells pairs. |
Copied to clipboard
| Challenge: | Existing work on summarization metrics and large language models has not explored fair abstractive summarizing. |
| Approach: | They propose four reference-free automatic metrics to measure the differences between target and source perspectives. |
| Outcome: | The proposed methods alleviate fair abstractive summarization on user-generated data. |
Copied to clipboard
| Challenge: | Existing dialogue state tracking models require plenty of labeled data, but collecting labels is expensive. |
| Approach: | They propose to use only 1% labeled data to train dialogue state tracking models . they encourage a model to have consistent latent distributions given a perturbed input . |
| Outcome: | The proposed self-supervised signals improve goal accuracy by 8.95% when only 1% labeled data is used on the MultiWOZ dataset. |
Copied to clipboard
| Challenge: | Existing text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies and contain strong layout and stylistic biases. |
| Approach: | They propose a dataset for long-form narrative summarization that uses human written summaries on three levels of difficulty. |
| Outcome: | The proposed dataset covers documents from the literature domain, such as novels, plays and stories, and includes highly abstractive, human written summaries on three levels of difficulty. |
Copied to clipboard
| Challenge: | Paraphrase generation has benefited from recent advances in the design of training objectives and model architectures, but previous studies focused on supervised methods that require a large amount of labeled data that is costly to collect. |
| Approach: | They propose a transfer learning approach that enables pre-trained language models to generate high-quality paraphrases in an unsupervised setting. |
| Outcome: | The proposed model performs state-of-the-art on the Quora Question Pair and ParaNMT datasets and is robust to domain shift between the two datasets. |
Copied to clipboard
| Challenge: | Neural networks lack the ability to reason about qualitative physics and cannot generalize to scenarios and tasks unseen during training. |
| Approach: | They propose a framework for reasoning about qualitative physics in natural language that generates interpretable descriptions of physical events. |
| Outcome: | The proposed framework generates explanations of how the physical simulation will causally evolve so that an agent or a human can reason about a solution using interpretable descriptions. |
Copied to clipboard
| Challenge: | Document interpretation and dialog understanding are the two major challenges for conversational machine reading. |
| Approach: | They propose a discourse-aware entailment reasoning network to strengthen the connection and enhance the understanding of document and dialog. |
| Outcome: | The proposed model improves document interpretation and dialog understanding on the ShARC benchmark. |
Copied to clipboard
| Challenge: | Vision-language models have broad competence that is difficult to evaluate . current evaluation benchmarks focus on only assessing one or a few capabilities . |
| Approach: | They perform a large-scale transfer learning experiment to discover latent VL skills from data. |
| Outcome: | The results suggest that factor analysis can identify reasonable yet surprising VL skill factors . the results contribute to the design of balanced and broad-coverage vision-language evaluation methods. |
Copied to clipboard
| Challenge: | Existing video large language models (LMMs) employ an impedance of thousands of frames to understand long videos. |
| Approach: | They propose a plug-and-play module integrated with VideoLLMs to facilitate efficient lengthy video perception. |
| Outcome: | The proposed module boosts the performance of open-source VideoLLMs and proprietary assistants on long-form video benchmarks. |
Copied to clipboard
| Challenge: | Existing methods that only address a fixed set of fields are difficult to use for different form types. |
| Approach: | They propose a value retrieval method with arbitrary queries for form-like documents . they propose 'docQueryNet' to predict target value based on understanding of layout and semantics of a form . |
| Outcome: | The proposed method outperforms existing methods on value retrieval . it improves document understanding on large-scale model pre-training by 17% . |
Copied to clipboard
| Challenge: | State-of-the-art models in NLP are opaque in terms of how they come to make predictions. |
| Approach: | They propose to release a benchmark to measure the quality of rationales extracted by models and how faithful these rationale are to human annotators. |
| Outcome: | The proposed benchmark will enable researchers to compare models and track progress on interpretable models for NLP. |
Copied to clipboard
| Challenge: | Experimental results show that state-of-the-art pretrained QA systems have limited zero-shot performance and tend to predict our questions as unanswerable. |
| Approach: | They propose a question-answering dataset that uses conversations as a knowledge source. |
| Outcome: | The proposed dataset provides a training and evaluation testbed to facilitate QA on conversations research. |
Copied to clipboard
| Challenge: | Existing methods for dialog state tracking are ontology-based and ontologie-free . however, it is not clear enough which slots are better handled by either of the two methods . |
| Approach: | They propose a dual-strategy model that integrates both ontology-based and ontological-free methods. |
| Outcome: | The proposed model outperforms the existing model on noisy and cleaner datasets. |
Copied to clipboard
| Challenge: | Pre-trained language models are often used to achieve state-of-the-art results . eval paper shows that generative language model can handle joint and multi-task settings . |
| Approach: | They propose to reformulate extraction and prediction tasks into a sequence generation task . they propose a generative language model with unidirectional attention that learns to accomplish the tasks via language generation . |
| Outcome: | The proposed model outperforms the state-of-the-art in few-shot and full-shot settings. |
Copied to clipboard
| Challenge: | Large Action Models (LAMs) face challenges due to the need for high-quality training data, especially for multi-steps tasks that involve planning, executing tool calls, and responding to feedback. |
| Approach: | They propose a framework for online exploration of agentic tasks with high-quality feedback . they use a dynamic task query generator and an extensive collection of tools to create a high-level feedback environment for LLM Agents. |
| Outcome: | The proposed framework achieves 49.3% performance improvement over baselines on toolbench and CRMArena. |
Copied to clipboard
| Challenge: | a current approach to solving NLP problems is to build a problem-specific dataset . current approaches do not allow for transforming tasks into textual entailment . |
| Approach: | They propose a pretrained textual entailment system that can generalize across domains . they argue that when is it worth transforming an NLP task into textual detailment? |
| Outcome: | The proposed model can generalize across domains with few examples, the authors argue . they show that it can be used for several downstream NLP tasks with limited annotations . |
Copied to clipboard
| Challenge: | Existing methods struggle to conduct deep searches and retrieve all necessary evidence. |
| Approach: | They propose a benchmark for evaluating deep search, a retrieval-augmented generation that requires source-aware, multi-hop reasoning over diverse, sparsed, but related sources. |
| Outcome: | The proposed benchmarks show that even the best-performing agentic RAG methods achieve an average performance score of 32.96 on the benchmark. |
Copied to clipboard
| Challenge: | Existing factual consistency benchmarks are inadequate to detect factual inconsistencies in LLMs. |
| Approach: | They propose a protocol for inconsistency detection benchmark creation and implement it in a 10-domain benchmark called SummEdits. |
| Outcome: | The proposed method is 20 times more cost-effective per sample and highly reproducible, as it estimates inter-annotator agreement at about 0.9. |
Copied to clipboard
| Challenge: | Existing pre-trained language models with self-attention encoder architectures are less useful in practice. |
| Approach: | They propose to use user and system tokens to model dialogue behavior during pre-training . they propose a contrastive objective function to simulate the response selection task . |
| Outcome: | The proposed model outperforms baseline models on four downstream tasks . it also has a few-shot ability that can mitigate the data scarcity problem . |
Copied to clipboard
| Challenge: | Existing approaches on semantic parsing suffer from exponential growth of logical form candidates and can hardly generalize to unseen data. |
| Approach: | They propose a unified semantic parser for question answering on KB and DB . they define the primitive as the essential element in their framework . |
| Outcome: | The proposed framework can predict logical forms by altering and composing top-ranked primitives with different operations. |
Copied to clipboard
| Challenge: | Existing workflow extraction methods for service agents are time-consuming and outdated, causing inconsistent and inconsistent results. |
| Approach: | They propose a framework for extracting and evaluating dialog workflows from historical interactions. |
| Outcome: | The proposed framework improves workflow extraction by 12.16% over baseline. |
Copied to clipboard
| Challenge: | Large language models have shown impressive performance in following natural language instructions to solve unseen tasks. |
| Approach: | They propose two strategies to help large language models better leverage task instructions . they propose to remove 60% of tokens from the task definitions while maintaining model performance . |
| Outcome: | The proposed approach achieves 4.2 Rouge-L improvement over 119 unseen test tasks. |
Copied to clipboard
| Challenge: | a recent study shows that multimodal models can reason across multiple modalities . a limited number of models are able to reason across a variety of inputs . |
| Approach: | They propose a dataset for contrastive cross-modal reasoning across four modalities . they use human annotations and a mixture-of-models round-trip-consistency filter . |
| Outcome: | a new model evaluates models on multiple modalities to determine which one best answers a natural language prompt . the model must select the one that best satisfies the query and then fine-tune it . state-of-the-art models still achieve only 56% accuracy overall and 42% in four-modal settings . |
Copied to clipboard
| Challenge: | Open-source vision-language models excel on simple question-answering tasks, but struggle with complex questions that require both perception and reasoning. |
| Approach: | They propose a family of vision-language models that have LeArned to Think wiTh vision spEcialists by offloading perception to state-of-the-art vision models. |
| Outcome: | The proposed model achieves 4-5% gains over baselines across 6 benchmarks covering both perception and reasoning abilities. |
Copied to clipboard
| Challenge: | a weakly-supervised approach is needed to verify factual consistency . auxiliary span extraction tasks are useful for verifying factual consistent summaries . |
| Approach: | They propose a weakly-supervised approach for verifying factual consistency . they transfer the model to summaries generated by several neural models . |
| Outcome: | The proposed approach outperforms models trained with strong supervision on source documents and human evaluations. |
Copied to clipboard
| Challenge: | Pre-trained multilingual language models show significant performance gains for zero-shot cross-lingual model transfer on a wide range of natural language understanding (NLU) tasks. |
| Approach: | They do cross-lingual evaluation using prompt tuning and compare it with fine-tuning . prompt tuning achieves much better cross-linguistic transfer than fine- tuning . |
| Outcome: | The results show that prompt tuning achieves better cross-lingual transfer than fine-tuning across datasets, with only 0.1% to 0.3% tuned parameters. |
Copied to clipboard
| Challenge: | Domain-adaptive post-training of large language models (LLMs) has emerged as promising approach for specialized domains such as medicine and finance. |
| Approach: | They propose a system to identify optimal adaptation criteria and training strategies for LLMs for the finance domain. |
| Outcome: | The proposed model achieves state-of-the-art performance across a wide range of financial tasks. |
Copied to clipboard
| Challenge: | Existing studies on multi-document summarization focus on collating information that all sources agree upon, but the task of summarizing diverse information remains underexplored. |
| Approach: | They propose a task of summarizing diverse information encountered in multiple news articles encompassing the same event using a dataset curated by a large language model. |
| Outcome: | The proposed task aims to summarize diverse information in multiple news articles encompassing the same event . the proposed task is difficult due to its limited coverage and verbosity biases . |
Copied to clipboard
| Challenge: | Existing evaluation frameworks for retrieval-augmented generation (RAG) systems focus on answerable queries, but ignore the importance of appropriately rejecting unanswerable requests. |
| Approach: | They propose a framework to evaluate whether retrieval-augmented generation systems handle unanswerable queries specific to a given knowledge base. |
| Outcome: | The proposed framework synthesizes diverse and challenging queries for any given knowledge base and evaluates them with unanswered ratio and acceptable ratio metrics. |
Copied to clipboard
| Challenge: | Existing work on few-shot intent classification without OOS has focused on the few-shot intent classification with out-of-scope intents. |
| Approach: | They propose to use BERT-style pairwise encoding to train a binary classifier that estimates the best matched training example for a user input. |
| Outcome: | The proposed approach achieves more stable and accurate in-domain and OOS detection accuracy than RoBERTa-based classifiers and embedding-based nearest neighbor approaches. |
Copied to clipboard
| Challenge: | Using pre-trained language models, we find out which model has the most informative representation for task-oriented dialogue tasks. |
| Approach: | They propose a supervised classifier probe and unsupervised mutual information probe to investigate the mutual dependence between a real clustering and a representation clustering. |
| Outcome: | The proposed model is a supervised classifier probe and unsupervised mutual information probe. |
Copied to clipboard
| Challenge: | Existing evaluation frameworks suffer from limitations such as static task benchmarks, limited scope, and inadequate integration with practical applications. |
| Approach: | They propose an open-source, Model Context Protocol-based evaluation framework specifically tailored for comprehensive and systematic assessment of LLM-powered agents. |
| Outcome: | The proposed framework uncovers nuanced performance patterns and identify domain-specific strengths and weaknesses, providing valuable insights beyond traditional binary success metrics. |
Copied to clipboard
| Challenge: | a new learning paradigm is proposed for NLP, which seeks supervision for solving a target task. |
| Approach: | They propose a new learning paradigm that uses textual instructions to learn new tasks . the main goal of machine learning algorithms lies in seeking supervision for solving a target task. |
| Outcome: | The proposed learning paradigm is based on a stream of more than 60 tasks . it makes full use of task instructions to improve forward-transfer and backward-transference . |
Copied to clipboard
| Challenge: | Prior work has shown that outcome-based reinforcement learning (RL) can incidentally elicit advanced reasoning behaviors such as self-correction, backtracking, and verification. |
| Approach: | They explicitly align large reasoning models with three meta-abilities: deduction, induction, and abduction, using automatically generated, self-verifiable tasks. |
| Outcome: | The proposed model aligns models with deduction, induction, and abduction meta-abilities using automatically generated, self-verifiable tasks. |
Copied to clipboard
| Challenge: | Existing approaches to dialogue state tracking are dependent on domain ontology and lack of sharing knowledge across domains. |
| Approach: | They propose a transferable dialogue state generator that generates dialogue states from utterances using copy mechanism. |
| Outcome: | Empirical results show that TRADE achieves state-of-the-art 48.62% joint goal accuracy for the five domains of MultiWOZ. |
Copied to clipboard
| Challenge: | Structured knowledge grounding (SKG) uses structured knowledge to complete user requests . since inputs and outputs of SKG tasks are heterogeneous, they have been studied separately . |
| Approach: | They propose a framework that unifies 21 SKG tasks into a text-to-text format . they use unifiedSKG to benchmark T5 with different sizes . |
| Outcome: | The proposed framework unifies 21 SKG tasks into a text-to-text format . it achieves state-of-the-art performance on almost all of the 21 tasks, the authors show . |
Copied to clipboard
| Challenge: | despite popularity of influence functions, their computational cost does not scale well with model and training data size. |
| Approach: | They propose a fast parallel variant that approximates the “influences” of training data-points for test predictions. |
| Outcome: | The proposed method achieves about 80X speedup while being highly correlated with the original influence values. |
Copied to clipboard
| Challenge: | Existing methods to debias word embeddings from human-generated corpora inherit strong gender bias . prior work has suggested removing gender component from pre-trained word embeds or compressing gender information into a few dimensions of the embeddable space . |
| Approach: | They propose a technique that purifies word embeddings against inferred gender subspaces . they propose to preserve distributional semantics of pre-trained word embeds while reducing gender bias . |
| Outcome: | The proposed technique preserves distributional semantics of pre-trained word embeddings while reducing gender bias to a larger degree than prior approaches. |
Copied to clipboard
| Challenge: | Existing methods to improve factual consistency of summarization models fail to remove entity errors if a suitable input entity replacement is not available or insert erroneous content. |
| Approach: | They propose to remove extrinsic entity errors, or entities not in the source, to improve consistency while retaining the summary’s essential information and form. |
| Outcome: | The proposed model improves factual consistency while maintaining ROUGE, improving entity precision by up to 30% on XSum, and can be applied on top of another post-editor, improving accuracy by 38%. |
Copied to clipboard
| Challenge: | Multi-hop reasoning is an effective approach for query answering over incomplete knowledge graphs (KGs). |
| Approach: | They propose to adopt a pretrained one-hop embedding model to estimate reward of unobserved facts and to force agents to explore diverse set of paths using randomly generated edge masks. |
| Outcome: | The proposed model reduces false negative supervision and counters spurious search trajectories by forcing the agent to explore a diverse set of paths using randomly generated edge masks. |
Copied to clipboard
| Challenge: | Existing natural language interfaces to databases are ambiguous or untranslatable . we present a robust, modular cross-domain text-to-SQL system . |
| Approach: | They propose a system that flags natural language input to which a SQL mapping cannot be immediately determined. |
| Outcome: | The proposed system can flag natural language input to which a SQL mapping cannot be determined. |
Copied to clipboard
| Challenge: | Modern news aggregators do the hard work of organizing the news, but choosing which source to read remains challenging. |
| Approach: | They propose a framework to help readers identify source differences and gain an understanding of news coverage diversity by generating questions with a diverse answer pool and reusing existing methods. |
| Outcome: | The proposed framework improves performance from current question generation methods by 5% and achieves 81% balanced accuracy on a realistic test set. |
Copied to clipboard
| Challenge: | Large language model (LLM)-based reasoning systems have recently achieved gold medal-level performance in the IMO 2025 competition . |
| Approach: | They propose a human-annotated step-level verification benchmark that measures step- level verifiers at the frontier. |
| Outcome: | The proposed benchmark outperforms closed-source models in step-level verification and the impact of scaling verifier compute. |
Copied to clipboard
| Challenge: | Recent work probes PLMs for the extent of factual knowledge through prompts . however, these methods do not consider symmetry of the task: object and subject prediction. |
| Approach: | They propose a continuous prompt-based method that leverages symmetry of the task by constructing symmetrical prompts for subject and object prediction. |
| Outcome: | The proposed method improves on a popular factual probing dataset on lAMA. |
Copied to clipboard
| Challenge: | Existing benchmarks for evaluating CRM agents on work-related tasks are limited due to data privacy concerns. |
| Approach: | They propose a benchmark to evaluate AI agents on real-world CRM tasks . they use 16 commonly used industrial objects with high interconnectivity to simulate real data distributions. |
| Outcome: | The new benchmark evaluates AI agents on real-world customer service tasks . it includes 16 commonly used industrial objects with high interconnectivity . the results highlight the need for enhanced agent capabilities in function-calling and rule-following . |
Copied to clipboard
| Challenge: | Current approaches to text summarization use advanced attention and copying mechanisms, multi-task and multi-reward training techniques. |
| Approach: | They evaluate datasets, evaluation metrics, and models for text summarization . they highlight three primary shortcomings: 1) datasets leave task underconstrained; 2) models overfit layout biases . |
| Outcome: | The current evaluation protocol is weakly correlated with human judgment and does not account for factual correctness. |
Copied to clipboard
| Challenge: | End-to-end neural networks excel at answering natural language questions but fail on complex ones . a proposed framework for question parsing and execution on textual QA is designed to combine the strengths of neural and symbolic methods. |
| Approach: | They propose a framework for question parsing and execution on textual QA . they parse questions into an intermediate representation and use deterministic rules to translate them . |
| Outcome: | The proposed framework outperforms existing methods in supervised, few-shot, and zero-shot settings while preserving its underlying reasoning process. |
Copied to clipboard
| Challenge: | Recent models infer latent representations of words or tokens with a transformer encoder, which is bottom-up and thus does not capture long-distance context well. |
| Approach: | They propose a method to infer latent representations of words or tokens in documents . they assume a hierarchical structure of a document where top-level captures long range dependency . |
| Outcome: | The proposed model can summarize an entire book and achieve competitive performance on a wide range of document summarization benchmarks. |
Copied to clipboard
| Challenge: | CoSQL is a corpus for building cross-domain, general-purpose database querying dialogue systems. |
| Approach: | They present a corpus for building cross-domain, general-purpose database querying dialogue systems . they use a Wizard-of-Oz collection of 3k turns plus 10k+ annotated SQL queries . |
| Outcome: | The proposed corpus is based on a Wizard-of-Oz dataset of 3k dialogues querying 200 complex DBs spanning 138 domains. |
Copied to clipboard
| Challenge: | Prior work on out-of-domain detection requires in-domain task labels and is limited to supervised classification scenarios. |
| Approach: | They propose a method to construct out-of-domain detectors efficiently using pre-trained transformers. |
| Outcome: | The proposed method greatly improves out-of-domain detection ability in a more general scenario. |
Copied to clipboard
| Challenge: | Existing studies for summarization evaluation exhibit low inter-annotator agreement or lack scale. |
| Approach: | They propose a modified summarization salience protocol based on fine-grained semantic units and a robust summarizing evaluation benchmark. |
| Outcome: | The proposed protocol is based on fine-grained semantic units and allows for high inter-annotator agreement. |
Copied to clipboard
| Challenge: | Autonomous agents powered by large language models (LLMs) have attracted significant research interest, but there are few standards for developing specialized models for agent tasks. |
| Approach: | They propose a series of large action models with dense and mixture-of-expert architectures that unifies, augments, and synthesizes diverse datasets to enhance agent generalizability and performance. |
| Outcome: | The proposed models outperform GPT-4, Claude-3, and many other models in terms of tool use and outperformed GPT-based models on multiple agent ability benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for synthesizing data for semantic parsing require handcrafted rules to synthesize new programs or utterance-program pairs. |
| Approach: | They propose to use a (non-neural) PCFG to model the composition of programs and a BART-based translation model to map a program to an utterance to learn a generative model from existing data. |
| Outcome: | The proposed model can be efficiently learned from existing data on benchmarks of GeoQuery and Spider. |
Copied to clipboard
| Challenge: | Existing evaluations of retrieval-augmented generation systems are limited . sub-question coverage measures how well a RAG system addresses different facets of a question. |
| Approach: | They propose a framework for evaluation based on sub-question coverage . they propose to decompose questions into sub-questions and classify them into three types . |
| Outcome: | The proposed evaluation framework measures how well a RAG system addresses different facets of a question. |
Copied to clipboard
| Challenge: | a global-local self-attentive dialogue state tracker estimates user goals and requests given the dialogue context . GLAD significantly improves tracking of rare states, compared to prior work . task-oriented dialogue systems can significantly reduce operating costs . |
| Approach: | They propose a global-local self-attentive dialogue state tracker which shares global-level modules with global-specific estimators for different types of dialogue states. |
| Outcome: | The proposed model outperforms previous models on the WoZ state tracking task by 3.9% and 4.8%. |
Copied to clipboard
| Challenge: | Existing conversational recommender systems focus on a single-shot approach to understand user preferences and provide recommendations. |
| Approach: | They propose a problem space for conversational agents that aim to provide both product recommendations and educational value through mixed-type mixed-initiative dialog. |
| Outcome: | The proposed framework can simulate salesbot and shopperbot agents and provide both product recommendations and educational value through mixed-type mixed-initiative dialog. |
Copied to clipboard
| Challenge: | Large language models have shown a powerful ability for text generation, but undesired behaviors such as toxicity and hallucinations can manifest. |
| Approach: | They propose to formalize text generation as a future-constrained generation problem to minimize undesirable behaviors and enforce faithfulness to instructions. |
| Outcome: | The proposed approach is effective across three tasks, including keyword-constrained generation, toxicity reduction, and factual correctness in question-answering. |
Copied to clipboard
| Challenge: | Empirical results indicate that we can effectively leverage language models for commonsense reasoning. |
| Approach: | They propose to use commonsense auto-generated explanations to train language models to generate explanations that can be used during training and inference in a commonsensense Auto-Generated Explanation framework. |
| Outcome: | Empirical results show that the proposed framework improves on the commonsenseQA task by 10%. |
Copied to clipboard
| Challenge: | Existing summarization systems produce generic summaries that are disconnected from users’ preferences and expectations. |
| Approach: | They propose a generic framework to control generated summaries through a set of keywords. |
| Outcome: | The proposed framework is comparable or better than strong pretrained systems on three domains of summarization datasets and five control tasks. |
Copied to clipboard
| Challenge: | Existing methods for evaluating progress in natural language generation tasks are expensive, difficult to reproduce, and non-reusable. |
| Approach: | They propose a new automatic evaluation method for NLG called Near-Negative Distinction that repurposes prior human annotations into NND tests. |
| Outcome: | The proposed method achieves higher correlation with human judgments than standard NLG evaluation metrics. |
Copied to clipboard
| Challenge: | Existing open-source frameworks like LangChain and LlamaIndex fail to integrate into daily workflows, resulting in limited daily usage for work. |
| Approach: | They propose a multi-agent library for scalable management and collaboration of AI agents on Slack. |
| Outcome: | The proposed framework offers instant AI integration into organizational workflows and facilitates scalable collaboration, allowing for effective communication and task orchestration. |
Copied to clipboard
| Challenge: | a lack of comprehensive studies on evaluation metrics for text summarization hinders progress . a new study aims to improve evaluation metrics that correlate with human judgments . |
| Approach: | They propose to re-evaluate automatic evaluation metrics and share a toolkit for evaluation . they hope to promote a more complete evaluation protocol for text summarization . |
| Outcome: | The proposed evaluation metrics are inconsistent with existing evaluation protocols. |
Copied to clipboard
| Challenge: | Ideal summarization models should generalize to novel summary-worthy content without remembering reference training summaries by rote. |
| Approach: | They propose to partition test set based on lexical similarity of reference test summaries with training summary to determine model competencies. |
| Outcome: | The proposed evaluation protocol improves generalization and generalization on novel test cases while maintaining average performance. |