Papers with metrics
Copied to clipboard
| Challenge: | a tutorial examines the future of work shaped by the interplay of large language models and humans . a series of tutorials examines challenges, opportunities, and ethical considerations in this dynamic landscape . |
| Approach: | This tutorial examines the future of work shaped by the interplay of LLMs and humans . it examines how LLM-based systems can augment human labor and enhance real-world tasks . |
| Outcome: | This tutorial examines the future of work shaped by the interplay of LLMs and humans . it examines challenges, opportunities, and ethical considerations in this dynamic landscape . |
Copied to clipboard
| Challenge: | Existing AVR benchmarks focus on single-step reasoning, emphasizing the end result but neglecting the multi-stage nature of reasoning process. |
| Approach: | They propose a multi-stage AVR benchmark based on RAVEN to assess reasoning across varying levels of complexity. |
| Outcome: | The proposed metric considers the correctness of intermediate steps in addition to the final outcomes. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have advanced capabilities but produce complex structured data. |
| Approach: | They propose a structure-aware fine-tuning method to bolster LLMs' performance by crafting format-specific instructions from the intended outputs. |
| Outcome: | The proposed method outperforms LLMs on all three formats and spans text tables, HTML, and LaTeX formats. |
Copied to clipboard
| Challenge: | Masked Language Models (MLMs) have shown strong performance on many NLP tasks. |
| Approach: | They evaluate masked language models for biomedical French on the task of clinical named entity recognition using gold-standard corpora. |
| Outcome: | The proposed model outperforms standard models on the task of clinical named entity recognition in biomedical French while remaining lighter than current models. |
Copied to clipboard
| Challenge: | This tutorial focuses on the evolution of voice-native LLMs . it reviews the adaptation of text LLM to audio, cross-modal alignment, and joint speech–text training . |
| Approach: | This tutorial examines the evolution of voice-native LLMs in conversational agents . it compares cascaded and voice-based LLM systems to end-to-end retrieval-and vision-grounded systems . |
| Outcome: | This tutorial examines the evolution of voice-native LLMs . it compares the performance of voice assistants to current open-domain agents . |
Copied to clipboard
| Challenge: | a tutorial will review the history of bias and fairness studies in machine learning and language processing . |
| Approach: | This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it presents recent community effort to quantify and mitigat bias in natural language processing models . |
| Outcome: | This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it aims to quantify and mitigate bias in natural language processing models for a wide spectrum of tasks . |
Copied to clipboard
| Challenge: | Automated headline generation systems have the potential to assist editors in finding interesting headlines to attract visitors or readers. |
| Approach: | They propose to use Bengali news article-headline pairings with auxiliary data to better model headline generation using pre-trained language models. |
| Outcome: | The proposed model improves on a Bengali news headline generation dataset by 3 to 10 percentage points over baselines. |
Copied to clipboard
| Challenge: | Existing work on vision and language navigation relies on navigation-related losses to establish the connection between vision and modalities, neglecting aspects of helping the navigation agent build a deep understanding of the visual environment. |
| Approach: | They propose to provide indirect supervision to the navigation agent through a hint generator that generates visual descriptions during navigation. |
| Outcome: | The proposed method improves the navigation performance and interpretability of the R2R and R4R datasets. |
Copied to clipboard
| Challenge: | Current Large Language Models (LLMs) benchmarks are often based on open-ended or close-ended QA evaluations, avoiding the requirement of human labor. |
| Approach: | They propose a multi-axis suite for healthcare LLM evaluation, exploring correlations between open and close benchmarks and metrics. |
| Outcome: | The proposed framework explores correlations between open and close benchmarks and metrics in the healthcare domain, with blind spots and overlaps in existing methodologies. |
Copied to clipboard
| Challenge: | Existing work on identifying the salient information in a text has used a limited representation of events that omits essential information. |
| Approach: | They propose a highly contextual model of event salience that uses a rich representation of events and integrates document-level information. |
| Outcome: | The proposed model improves on an event salience dataset by 2-4% on standard metrics and addresses flaws in existing evaluation methodologies. |
Copied to clipboard
| Challenge: | Conventionally, keyword decision-making in sponsored search advertising relies on deep generation-based methods. |
| Approach: | They propose an LLM agent-based method that dynamically monitors KPI changes and adapts keyword generation in real-time. |
| Outcome: | The proposed method shows significant improvements across various metrics and emphasizes the importance of each component. |
Copied to clipboard
| Challenge: | Language models (LMs) have exhibited impressive abilities in generating code from natural language requirements. |
| Approach: | They propose to introduce various metrics with inter-code similarity to evaluate the diversity of generated code by comparing model-generated solutions with human-written ones. |
| Outcome: | The proposed method leverages LMs’ capabilities in code understanding and reasoning, resulting in a set of metrics that represent the number of algorithms in model-generated solutions. |
Copied to clipboard
| Challenge: | Recent advances in generative spoken language modeling have produced models that produce speech in a wide range of voices, prosody and recording conditions. |
| Approach: | They propose acoustic diversity metrics that measure voice, gender, emotion, accent, background noise and a priori known diversity preferences for each facet. |
| Outcome: | The proposed metrics show that they achieve stronger agreement with diversity than baselines. |
Copied to clipboard
| Challenge: | Evaluation is a key part of machine learning, yet there is neo-tooling to support it . auxiliary techniques such as testing for significance, measuring statistical power, and auxiliary methods are not available in ML. |
| Approach: | They propose a set of tools to facilitate the evaluation of models and datasets in machine learning . they propose 'evaluation on the Hub' platform that enables large-scale evaluation of over 75,000 models . |
| Outcome: | The proposed tools can be used to evaluate models and datasets on the Hugging Face Hub. |
Copied to clipboard
| Challenge: | Existing evaluation tools for general Visual Question Answering (VQA) systems are limited to answering accuracy, but they can be used to evaluate performance in real-world scenarios. |
| Approach: | They propose a browser-based benchmarking tool with an API for easy integration of new models and datasets to keep up with the fast-changing landscape of VQA. |
| Outcome: | The proposed tool tests generalization capabilities of models across multiple datasets and includes metrics that measure biases and uncertainty to further explain model behavior. |
Copied to clipboard
| Challenge: | Current evaluation methods focus on one dataset, e.g., Newstest dataset in each year’s WMT Metrics Shared Task. |
| Approach: | They propose to use a single dataset to evaluate the performance of automatic translation metrics. |
| Outcome: | The results show that the rankings of metrics vary when the evaluation is conducted on different datasets. |
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) is a popular NLG evaluation metric . however, it is difficult to judge where exactly such a metric fails . |
| Approach: | They propose a checklist for NLG evaluation metrics that focus on meaning by organizing them around meaning-relevant linguistic phenomena. |
| Outcome: | The proposed metric GraCo computes lexical cohesion graphs over AMR concepts. |
Copied to clipboard
| Challenge: | Neural models have shown significant progress on data-to-text generation tasks . data- to-text models generate descriptive texts from non-linguistic structured data . |
| Approach: | They propose a new data-to-text generation model which learns content selection and summary generation in an end-to end fashion. |
| Outcome: | The proposed model outperforms current state-of-the-art models on content selection precision and content ordering metrics. |
Copied to clipboard
| Challenge: | Existing n-gram based QA metrics have a number of drawbacks and are not suitable for all extractive tasks. |
| Approach: | They propose to use BERTScore to evaluate translation for question answering (QA) they also explore whether existing n-gram based metrics are suitable for generative QA . |
| Outcome: | The proposed BERTScore metric fails to provide stronger correlation with human judgements . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) generate text that is stereotypical or not representative of the viewpoints and values of historically marginalized demographic groups. |
| Approach: | They propose to use data from the Olympic Games to investigate gender bias in large language models. |
| Outcome: | The proposed model consistently biased against women when the gender is ambiguous in the prompt, revealing pervasive gender bias in LLMs in the context of athletics. |
Copied to clipboard
| Challenge: | Existing table-to-text generation benchmarks have some limitations, such as E2E and ToTTo focusing on singlesentence generation tasks. |
| Approach: | They propose a new table-to-text generation dataset called TaKG that uses a set of knowledge graphs to enhance table input. |
| Outcome: | The proposed model outperforms existing models for short-text generation tasks and shows reliable performance on long-text generated across a variety of metrics. |
Copied to clipboard
| Challenge: | MGTD methods are needed in many areas, such as prevention of disinformation spreading, plagiarism, impersonation and identity theft. |
| Approach: | They propose a framework for machine-generated text detection that integrates custom methods and evaluation datasets into existing frameworks. |
| Outcome: | The proposed framework simplifies the benchmarking of machine-generated text detection methods by easy integration of custom (new) methods and evaluation datasets. |
Copied to clipboard
| Challenge: | Existing benchmarks are narrow and simply compute overall task success. |
| Approach: | They propose a framework where both benchmarks and metrics are modular and easily extensible through well documented and easy-to-use APIs. |
| Outcome: | The proposed framework can track agent progress on two use cases and identify common failure points and refine the agent architecture to obtain a significant performance increase. |
Copied to clipboard
| Challenge: | Recent work in cognitive neuroscience has introduced models for predicting distributional word meaning representations from brain imaging data. |
| Approach: | They propose to use several alternative measures to evaluate the predicted distributional space against a corpus-derived distributional spatial space. |
| Outcome: | The proposed model performs poorly on the most common metrics, while still delivering promising results. |
Copied to clipboard
| Challenge: | Existing citation recommendation systems rely on information of query documents such as author names and publication venue. |
| Approach: | They propose a content-based method for recommending citations in academic paper drafts . they embed a given query document into a vector space and use its nearest neighbors as candidates . |
| Outcome: | The proposed method outperforms published methods on PubMed and DBLP datasets without metadata. |
Copied to clipboard
| Challenge: | Existing methods for object navigation are limited to household datasets with close-set objects, and they lack the ability to generalize to new environments in a zero-shot manner. |
| Approach: | They propose a framework that leverages reasoning abilities of large vision language models to extract proposed objects from natural language instructions that meet the user’s demand. |
| Outcome: | The proposed framework surpasses baselines on all metrics and can be used in a HM3D ObjectNav benchmark. |
Copied to clipboard
| Challenge: | MolT5 pretrains models on unlabeled natural language text and molecule strings . bringing a new drug to market can cost over a billion dollars and take over ten years . |
| Approach: | They propose a self-supervised learning framework for pretraining models on unlabeled natural language text and molecule strings. |
| Outcome: | The proposed framework pretrains models on unlabeled natural language text and molecule strings, and it generates high quality outputs. |
Copied to clipboard
| Challenge: | BU-NEmo dataset extends from 320 to 1,297 news headline and lead image pairings and collects 38,910 annotations in a crowdsourcing experiment. |
| Approach: | They extend the U.S. gun violence news-to-emotions dataset from 320 to 1,297 news headline and lead image pairings and collect annotations in a crowdsourcing experiment. |
| Outcome: | The proposed models outperform baseline models on the NEmo+ dataset by large margins across several metrics. |
Copied to clipboard
| Challenge: | Existing models that generate free-text explanations for tasks are limited by human-written explanations. |
| Approach: | They propose to use a standardized collection of natural language prompts to create a model that generates free-text explanations for tasks. |
| Outcome: | The proposed model can predict task labels and generate free-text explanations for predictions . plausibility of human explanations is 76%, while human explanation is 51% . |
Copied to clipboard
| Challenge: | Existing evaluation metrics for dialog state tracking are limited for belief states accumulated as dialog proceeds . relative slot accuracy allows intuitive evaluation by assigning relative scores according to the turn of each dialog . |
| Approach: | They propose to use relative slot accuracy to complement existing evaluation metrics . joint goal accuracy and slot accuracy are used to evaluate accumulated belief states . |
| Outcome: | The proposed metrics focus on "penalizing states that fail to predict," not "reward for well-predicted states" the proposed metrics do not depend on the number of predefined slots, and allow intuitive evaluation . |
Copied to clipboard
| Challenge: | Existing de-identification methods suffer from recall errors, limited generalization, and inefficiencies, limiting their real-world applicability. |
| Approach: | They propose a multi-modal framework for de-identifying electronic health records using a retrieval-based entity relexicalization approach. |
| Outcome: | The proposed framework achieves competitive performance while optimizing token usage to reduce LLM costs. |
Copied to clipboard
| Challenge: | Existing studies on LLM factuality evaluation have not investigated the reliability of static evaluation benchmarks. |
| Approach: | They examine five popular factuality benchmarks and eight LLMs released over different years to assess their reliability. |
| Outcome: | The proposed method compared five popular factuality benchmarks and eight LLMs released over different years. |
Copied to clipboard
| Challenge: | Minimum Bayes risk (MBRS) decoding is a decision rule of text generation tasks that outperforms conventional maximum a posteriori (MAP) decoders by selecting high-quality outputs based on quality or preference rather than probability. |
| Approach: | They propose to use minimum bayes risk (MBRS) decoding to determine outputs based on quality rather than probability. |
| Outcome: | MBRS is an MIT-licensed open-source project with a focus on speed, reproducibility, and extensibility. |
Copied to clipboard
| Challenge: | Statistical fairness stipulates equivalent outcomes for all protected groups, whereas causal fairness prescribes that a model makes the same prediction for an individual regardless of their protected characteristics. |
| Approach: | They propose to use statistical and causal debiasing methods to reduce gender bias in NLP models. |
| Outcome: | The proposed methods reduce gender bias measured by the targeted metric, but not on other bias metrics. |
Copied to clipboard
| Challenge: | Neural networks are notoriously hard to interpret and slightly mysterious to researchers and practitioners alike. |
| Approach: | They formalize hyperparameter sensitivity using two metrics: similarity-based sensitivity and performance-based-sensitivity. |
| Outcome: | The transformer is more sensitive to hyperparameters according to both metrics, but not batch size . large models, multilinguality of NLP models and tasks make hyperparametric tuning more expensive . |
Copied to clipboard
| Challenge: | Existing studies on pre-trained vision-language models have focused on measuring biases and stereotypes in a single modality. |
| Approach: | They extend a recently released stereotypical bias dataset into a vision-language probing dataset called VLStereoSet to measure stereotypical biased vision-linguistic models. |
| Outcome: | The proposed probing task measures stereotypical bias in vision-language models and its intra-modal and inter-modal biases. |
Copied to clipboard
| Challenge: | Recent advances in text generation systems produce fluent, coherent, relevant, and factually correct text. |
| Approach: | They propose a metaevaluation framework for evaluating factuality evaluation metrics . they propose five necessary conditions to evaluate factual metrics on diagnostic factuity data . |
| Outcome: | The proposed framework provides robust evaluation that is extensible to multiple types of factual consistency and standard generation metrics, including QA metrics. |
Copied to clipboard
| Challenge: | Word embeddings have been used to quantify biases in texts for years, but their statistical properties and advantages have not been studied. |
| Approach: | They propose to use PMI-based metric to quantify bias in corpora by conditional probabilities and odds ratio to approximate it. |
| Outcome: | The proposed measure can be approximated by an odds ratio, which makes statistical inferences cost-effective and meaningful. |
Copied to clipboard
| Challenge: | Image captioning is a core task in multimodal NLP, where the aim is to automatically describe the content of an image in natural language. |
| Approach: | They propose to use syntactic tags and tokens to improve caption generalization . they also propose to model the syntakic structure of a caption to improve generalization. |
| Outcome: | The proposed models improve generalization and performance on standard metrics while requiring syntactic and semantic knowledge of the language. |
Copied to clipboard
| Challenge: | Existing automated question generation methods focus on unstructured text and lack agenda and background documents as context. |
| Approach: | They propose to leverage large language models for CourtQG by fine-tuning them on two auxiliary tasks, agenda explanation and question type prediction. |
| Outcome: | The proposed method generates better questions according to standard metrics when compared to several baselines. |
Copied to clipboard
| Challenge: | aggregating results over incomparable metrics and scenarios makes conclusions and take-away messages less reliable . |
| Approach: | They propose a task-agnostic toolkit that combines the effect of a treatment on multiple tasks into one statistical evaluation, allowing comparison of metrics and computation of an overall summary effect. |
| Outcome: | The proposed toolkit produces publication-ready forest plots that enable clear communication of evaluation results over multiple tasks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly being applied to causal reasoning tasks. |
| Approach: | They propose a symbolic verification framework that checks whether LLM-generated causal expressions are derivable from a given causal graph using do-calculus and probability theory. |
| Outcome: | The proposed framework can recover correct answers that would otherwise be marked incorrect due to superficial differences. |
Copied to clipboard
| Challenge: | elucidates the dangerous current state of style transfer auto-evaluation research. |
| Approach: | They propose ways to aggregate the three metrics into one evaluator. |
| Outcome: | The proposed method could be used to aggregate the three metrics into one evaluator. |
Copied to clipboard
| Challenge: | Existing query rewriting models ignore user history behaviors and consider only the instant search query, which is often a short string offering limited information about the true shopping intent. |
| Approach: | They propose an end-to-end context-aware query rewriting model that takes search context into account and builds a session graph using the history search queries and their contained words. |
| Outcome: | The proposed model outperforms state-of-the-art models under various metrics. |
Copied to clipboard
| Challenge: | Existing methods for generating concise and coherent summaries may include unintended text with hallucinations, causing computational costs. |
| Approach: | They propose a model-agnostic framework to post-process medical hallucinations . MEDAL integrates with any medical summarization model, requiring no additional computational overhead . |
| Outcome: | MEDAL can post-process medical hallucinations without additional computational overhead. |
Copied to clipboard
| Challenge: | In machine translation evaluation, metric performance is assessed based on agreement with human judgments. |
| Approach: | They incorporate human baselines into the MT meta-evaluation to gain a clearer understanding of metric performance and establish an upper bound. |
| Outcome: | The results suggest human parity, but there are several reasons to caution . |
Copied to clipboard
| Challenge: | Existing AMR metrics are inefficient and struggle to capture semantic similarity . Existing metrics are not efficient and lack a systematic evaluation benchmark . |
| Approach: | They propose a new AMR similarity metric, rematch, which matches graphs structurally and semantically to each other. |
| Outcome: | The proposed metric is five times faster than the next most efficient metric. |
Copied to clipboard
| Challenge: | Existing methods for temporal reasoning have been used for a number of applications, but their potential for tempor reasoning over event graphs has not been explored. |
| Approach: | They propose to use large-scale pre-trained language models to generate an event-level temporal graph from a document using existing IE/NLP tools. |
| Outcome: | The proposed method outperforms the closest existing method on several metrics on a hand-labeled, out-of-domain corpus. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) excel in various tasks, but their evaluation, especially in languages beyond the top 20, remains inadequate due to existing benchmarks and metrics limitations. |
| Approach: | They propose to use Large Language Models as evaluators to rank or score other models’ outputs by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages. |
| Outcome: | The proposed evaluation methods can be used to improve multilingual evaluation by calibrating them against 20K human judgments across three text-generation tasks, five metrics, and eight languages. |
Copied to clipboard
| Challenge: | Recent advances in pretrained language models have shown promising results on commonsense reasoning benchmark datasets. |
| Approach: | They propose a commonsense reasoning benchmark dataset with 4k sentence pairs . they propose 'gamified' model-in-the-loop setup to incentivize challenging samples . |
| Outcome: | The proposed benchmarks show that the proposed model achieves 71% standard accuracy and 51% pairwise accuracy, well below human performance. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) face limitations due to outdated knowledge, hallucinations, and poor reasoning in complex contexts. |
| Approach: | They propose a Hybrid Parameter-Adaptive RAG system for the AI legal domain with NYC Local Law 144 as the test case. |
| Outcome: | The proposed system improves retrieval accuracy, response fidelity, and contextual precision on NYC Local Law 144 . Empirical evidence indicates that many AI tools overstate their ability to prevent hallucinations in legal and policy contexts. |
Copied to clipboard
| Challenge: | Conditional language models generate unfaithful output that is not supported by their input . this jeopardizes trust in real-world applications, raising a need for automatic faithfulness metrics. |
| Approach: | They propose to augment conditional language models with robust inference procedures to improve faithfulness. |
| Outcome: | The proposed approach outperforms existing models on the TRUE benchmark. |
Copied to clipboard
| Challenge: | Text-to-speech (TTS) synthesis has seen significant advancements in recent years. |
| Approach: | They propose to use PhoAudiobook to curated 941 hours of high-quality audio for Vietnamese text-to-speech models. |
| Outcome: | The proposed model improves on VALL-E, VoiceCraft, and XTTS-V2 models, highlighting their robustness in handling diverse linguistic contexts. |
Copied to clipboard
| Challenge: | Out-of-distribution (OOD) detection is a fundamental task vexing real-world applications . fine-tuning based methods require storing fine- tuned models for each scenario . |
| Approach: | They propose an unsupervised prefix-tuning based OOD detection framework called PTO . they propose to take advantage of optional training data labels and targeted OOD data . |
| Outcome: | The proposed framework performs better than existing methods under a wide range of metrics, detection settings, and OOD types. |
Copied to clipboard
| Challenge: | Sparse Autoencoders (SAEs) are a promising tool for disentangling FM representations, but they struggle to capture rare, yet crucial concepts in the data. |
| Approach: | They propose a technique to train Sparse Autoencoders to illuminate elusive dark matter features by focusing on specific subdomains. |
| Outcome: | The proposed method achieves 12.5% better classification accuracy than general-purpose SAEs when applied to remove spurious gender information. |
Copied to clipboard
| Challenge: | Knowledge Graphs (KGs) are a powerful tool for capturing structured representations of the world. |
| Approach: | They propose a scalable method for generating up-to-date and configurable conversational KGQA datasets that adheres to human interaction configurations and operates at a significantly larger scale. |
| Outcome: | Qualitative psychometric analyses show that ConvKGYarn produces high-quality data comparable to popular conversational KGQA datasets across various metrics. |
Copied to clipboard
| Challenge: | Abstractive summarization is a core application in contact centers, where Large Language Models generate millions of summaries of call transcripts daily. |
| Approach: | They propose a framework that uses an LLM as a zero-shot classifier to derive categorical distributions for each bias dimension in a pair of transcripts and its summary. |
| Outcome: | The proposed framework identifies and quantifies 15 operational bias dimensions and measures them using two metrics: Fidelity Gap and Coverage. |
Copied to clipboard
| Challenge: | Recent work suggests strategies to increase inference efficiency with LLMs . however, these strategies may inadvertently lead to some side-effects. |
| Approach: | They propose to optimize inference acceleration strategies such as quantization, pruning, and caching to reduce inference cost and latency while maintaining predictive performance. |
| Outcome: | The proposed strategies reduce cost and latency while maintaining predictive performance while preserving the model size. |
Copied to clipboard
| Challenge: | Existing automatic evaluation metrics are based on procedures that diverge from human evaluation. |
| Approach: | They propose to aggregate automatic evaluation metrics to bridge this gap . they propose to use edit-based metrics, -gram based metrics and sentence-level metrics to find the best ranking system. |
| Outcome: | The proposed method outperforms existing metrics on the SEEDA benchmark and improves edit-based metrics, -gram based metrics and sentence-level metrics. |
Copied to clipboard
| Challenge: | Existing explanations address the contrastive aspect of explanations but their extension to textual data is under-explored and there is little investigation on their vulnerabilities and limitations. |
| Approach: | They propose a novel evaluation scheme inspired by the faithfulness of explanations by extending the computation of three metrics to textual data and benchmarking POLYJUICE and MiCE on suggested metrics. |
| Outcome: | The proposed methods demonstrate that the connectedness of counterfactuals to their original counterparts is not obvious in both models. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have made significant progress in integrating safety and knowledge alignment, but excessive focus on safety alignment can lead to unintended hallucinations. |
| Approach: | They propose a "safety-priming" method to generate synthetic safety data and overcome safety bottlenecks. |
| Outcome: | The proposed framework generates synthetic safety data and overcomes safety bottlenecks. |
Copied to clipboard
| Challenge: | Large Language Models encode substantial factual knowledge, yet measuring and systematizing it remains challenging. |
| Approach: | They systematically analyze LLM knowledge materialization using miniGPTKBs . they find high termination rates, though model-dependent, and mixed reproducibility . |
| Outcome: | The proposed model can reliably surface core knowledge, but it has limitations. |
Copied to clipboard
| Challenge: | Reliably evaluating Machine Translation (MT) through automated metrics is a long-standing problem. |
| Approach: | They propose to use MT models to generate multiple diverse translations and use them as surrogates to reference translations to obtain a quantification of translation variability. |
| Outcome: | The proposed approach improves correlation with human judgements of quality by 15%. |
Copied to clipboard
| Challenge: | Existing methods for generating text are unsupervised and require supervision. |
| Approach: | They propose an unsupervised method that uses two off-the-shelf pretrained LMs in opposite directions to apply them to non-sequential tasks. |
| Outcome: | The proposed method outperforms strong unsupervised baselines on paraphrasing and abductive text infilling. |
Copied to clipboard
| Challenge: | et al., 2018: translation errors due to the lack of extra-sentential context are becoming more and more noticeable among otherwise adequate translations. |
| Approach: | They propose a context-aware translation model that uses sentence-level data to identify inconsistencies . standard metrics are not sensitive to improvements in consistency in document-level translations . |
| Outcome: | The proposed model shows major gains over baseline without sacrificing performance . standard metrics are not sensitive to improvements in document-level translations . |
Copied to clipboard
| Challenge: | Existing methods to associate geographic information in text with coordinates are limited by lexical features and cartesian coordinates. |
| Approach: | They propose a geocoder that exploits implicit lexical clues to associate coordinates with text . they propose encoding of geographic metadata to generate two distinct views of the same text. |
| Outcome: | The proposed method improves state-of-the-art results on three datasets and an open-source dataset for disease outbreaks and epidemics. |
Copied to clipboard
| Challenge: | Using standard metrics in the presence of poor labels masks label and model quality . evaluation techniques accounting for unreliable labels reveal important flaws, including spurious correlations and nonrandom racial biases . |
| Approach: | They analyze human labels, GPT model ratings, and transformer encoder model ratings . they show that standard metrics in the presence of poor labels mask label and model quality . |
| Outcome: | The proposed methods mask label and model quality even in the presence of poor models. |
Copied to clipboard
| Challenge: | generating aspect-specific and general opinion summaries is challenging due to the lack of annotated data. |
| Approach: | They propose two unsupervised approaches to generate aspect-specific and general opinion summaries by training on synthetic datasets constructed with aspect-related review contents. |
| Outcome: | The proposed method outperforms existing methods on space and Oposum+ and on other metrics. |
Copied to clipboard
| Challenge: | a recent survey of bias in natural language processing found that a coreference system makes more errors in an anti-stereotypical coreferent than in a pro-sterereotype one. |
| Approach: | They compare intrinsic and extrinsic bias metrics across hundreds of trained models . they urge researchers to focus on extrindic measures of bias, not easy to measure . |
| Outcome: | a new intrinsic metric and an annotated test set on gender bias in hate speech are tested . authors urge researchers to focus on extrinsic measures of bias, and to make them more feasible . |
Copied to clipboard
| Challenge: | Recent studies show that the energy requirements of current NLP models are growing at a rapid, unsustainable pace. |
| Approach: | They investigate ways to measure energy usage and different hardware settings that can be tuned to reduce energy consumption for training and inference for language models. |
| Outcome: | The proposed techniques can reduce energy consumption for training and inference for language models. |
Copied to clipboard
| Challenge: | Chain-of-Thought (CoT) prompting can dramatically improve the multi-step reasoning abilities of large language models (LLMs). |
| Approach: | They propose to use Chain-of-Thought (CoT) prompting to encourage the LLM to generate intermediate rationales for solving a problem by providing a series of reasoning steps in the demonstrations. |
| Outcome: | The proposed model can generate coherent lines of reasoning even with invalid demonstrations while still generating coherent lines during inference. |
Copied to clipboard
| Challenge: | Existing research assesses LLMs’ values by analyzing their stated inclinations . a framework to evaluate the alignment between stated values and value-informed actions is lacking . |
| Approach: | They propose a framework to evaluate the alignment between LLMs’ stated values and their value-informed actions. |
| Outcome: | The proposed framework shows significant misalignment between LLM-generated values and their actions . misaligned values have shown real-world risks, such as amplifying stereotypes and reinforcing bias algorithms in hiring. |
Copied to clipboard
| Challenge: | Contemporary statistical models trade off interpretability and simplicity for powerful parameterizations and inductive biases, enabling impressive performance. |
| Approach: | They examine three recent models and find they are not yet reliable . they also formulate recommendations for practitioners and researchers . |
| Outcome: | The proposed models are not as reliable as previously assumed, the authors argue . their findings suggest that they are needed for improving models and training setups . |
Copied to clipboard
| Challenge: | a new schema for NLP knowledge about tasks, datasets and metrics is proposed. |
| Approach: | They propose a new schema that represents knowledge about tasks, datasets and metrics in the NLP domain. |
| Outcome: | The proposed framework can be automatically built into scientific leaderboards . the proposed system achieves reasonable results for all relation types on this small-scale graph . |
Copied to clipboard
| Challenge: | Dialogue systems that generate factually incorrect responses are often unfitful and hallucinate factuality invalid. |
| Approach: | They propose a method to improve faithfulness and reduce hallucination of neural dialogue systems to known facts supplied by a Knowledge Graph. |
| Outcome: | The proposed approach improves faithfulness and reduces hallucination of dialogue systems to known facts . it leverages a token-level fact critic to identify plausible sources of hallucinism . |
Copied to clipboard
| Challenge: | Existing conditional text generation models produce unfaithful and unfaithed summaries . current models accomplish a high level of fluency and coherence . |
| Approach: | They propose to use pretrained models for document summarization to better understand hallucinations . they find that textual entailment measures better correlate with faithfulness . |
| Outcome: | The proposed models generate faithful and factual summaries as evaluated by humans. |
Copied to clipboard
| Challenge: | Evaluating the performance of LLMs in multi-turn interactions presents significant challenges due to the complexity and variability of user behavior. |
| Approach: | They propose a benchmark framework for assessing LLMs’ function-calling capabilities in multi-turn dialogues. |
| Outcome: | The proposed framework is based on a dataset derived from popular mobile apps and anonymized user logs. |
Copied to clipboard
| Challenge: | Large language models (LLMs) trained over corpora risk memorizing sensitive, copyrighted, or toxic content. |
| Approach: | They propose a framework that removes targeted data while preserving model utility. |
| Outcome: | The proposed framework resists membership inference attacks, minimizes impact on retained data, and maintains robustness across diverse scenarios. |
Copied to clipboard
| Challenge: | Large language models (LLMs) increasingly power car assistants, but evaluating response quality remains a challenge. |
| Approach: | They propose a framework that uses large language models as evaluators to compare assistant responses against ground-truth counterparts. |
| Outcome: | The proposed framework compares assistant responses against ground-truth counterparts, assessing coverage, correctness, and other dimensions of answer quality. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for NRG models can't measure semantic relevance and diversity of generated results. |
| Approach: | They propose a large-scale domain-specific conversational corpus with preprocessing and cleansing procedures for model training and a testing set for measuring the diversity of generated results. |
| Outcome: | The proposed corpus can be taken as a new benchmark dataset for the NRG task. |
Copied to clipboard
| Challenge: | Existing pre-trained language models are not fully considered for societal biases . pre-training models can be useful for many NLP tasks, but they can be harmful when used at scale. |
| Approach: | They investigate gender and racial bias across pre-trained language models . they evaluate bias within pre-trainers using three metrics: WEAT, sequence likelihood, and pronoun ranking. |
| Outcome: | The proposed model fails to detect gender and racial biases in pre-trained models . the model is ineffective when word embedding, demonstrating the need for more robust bias testing in transformers. |
Copied to clipboard
| Challenge: | Existing methods for attribution of knowledge in large language models struggle to operate at neuron level due to computational constraints. |
| Approach: | They propose a static method for pinpointing significant neurons using three metrics . they also propose identifying "query neurons" which activate these "value neurons" |
| Outcome: | The proposed method shows superior performance across three metrics compared to seven other methods . it analyzes six types of knowledge across attention and feed-forward network layers . |
Copied to clipboard
| Challenge: | Existing evaluation frameworks do not assess why a text is deemed hateful . authors present a new metric to evaluate the reasoning quality of model explanations . |
| Approach: | They propose a metric suite to evaluate the reasoning quality of model explanations. |
| Outcome: | The proposed metric validates it as a practical tool for trustworthy and transparent moderation on six diverse hate speech datasets. |
Copied to clipboard
| Challenge: | Existing approaches to hierarchical text classification focus on parent-child relationships . however, some texts with a category hierarchy also have latent relevancy among labels in the same level of the hierarchy. |
| Approach: | They propose a method to analyze latent relevancy of peer labels and a sample importance learning method to ameliorate the side effects. |
| Outcome: | The proposed method improves the latent relevancy of peer labels on standard datasets. |
Copied to clipboard
| Challenge: | Existing lexical resources for semantic annotation of synonyms are lacking in computational language processing. |
| Approach: | They describe a bilingual lexical resource being built to investigate verbal synonymy in bilingual context and relate semantic roles common to one synonym class to verb arguments. |
| Outcome: | The proposed resource is based on English and Czech WordNet, FrameNet, PropBank, VerbNet (SemLink), and valency lexicons for Czech and English (PDT-Vallex, Vallex, and EngValleX). |
Copied to clipboard
| Challenge: | Existing evaluation metrics for paraphrase generation are not designed for the task, but adopted from other evaluation tasks. |
| Approach: | They propose a new evaluation metric for paraphrase generation that uses reference-based and reference-free metrics. |
| Outcome: | The proposed evaluation metric outperforms existing metrics and is more reliable than reference-based metrics. |
Copied to clipboard
| Challenge: | Existing automatic dialogue coherence evaluation metrics are expensive and high-latency, which cannot meet the requirements of a dialogue system. |
| Approach: | They propose a framework to train a quantifiable dialogue coherence metric that can reflect actual human rating standards. |
| Outcome: | Experimental results show that the model trained by QuantiDCE presents stronger correlations with human judgements than the other state-of-the-art metrics. |
Copied to clipboard
| Challenge: | Existing dialogue models struggle to interpret context accurately due to irrelevant or misclassified knowledge, limiting their effectiveness in real-world scenarios. |
| Approach: | They propose a framework that dynamically filters relevant commonsense knowledge and extracts personalized information to improve empathetic dialogue generation. |
| Outcome: | The proposed framework outperforms existing models in coherence, emotional understanding, and response relevance on the ESConv dataset. |
Copied to clipboard
| Challenge: | Recent generative language models have shown promise in abstractive summarization tasks. |
| Approach: | They propose to use Fr echet embedding distance and angular embeddable similarity to evaluate the performance of generative language models in abstractive summarization tasks. |
| Outcome: | The proposed metric shows close relation with human judgments and has overall better correlations with them. |
Copied to clipboard
| Challenge: | Existing state-of-the-art VLN agents do not generalize well for long navigation tasks. |
| Approach: | They propose a VLN agent that is learned to navigate by decomposing long instructions into shorter ones and completing them sequentially. |
| Outcome: | The proposed agent can follow long instructions better than existing ones, but it does not generalize well. |
Copied to clipboard
| Challenge: | Extensive experiments with seven Large Language Models reveal their varying behaviors. |
| Approach: | They investigate the behaviors of Large Language Models when faced with conflicting prompts versus their internal memory. |
| Outcome: | Extensive experiments with seven LLMs reveal their varying behaviors. |
Copied to clipboard
| Challenge: | a recent study has shown that fonts with a large number of missing glyphs are difficult to model due to the relative sparsity of most fonts. |
| Approach: | They propose a deep generative model that performs typography analysis and font reconstruction by learning disentangled manifolds of both font style and character shape. |
| Outcome: | The proposed model scales up the number of character types we can model compared to previous methods . it can generalize to characters that were not observed during training time, and it compares favorably to other models . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have highlighted the need for effective unlearning mechanisms to comply with data regulations and ethical AI practices. |
| Approach: | They propose a second-order optimization-based LLM unlearning framework which extends the static, one-shot model update using influence unlearning to a dynamic, iterative unlearning process. |
| Outcome: | The proposed framework outperforms first-order methods across unlearning tasks, models, and metrics. |
Copied to clipboard
| Challenge: | Embedding-based models are increasingly needed for domain-specific evaluation datasets. |
| Approach: | They propose a protocol for the construction of a relatedness-based evaluation dataset based on adaptive pairwise comparisons and appropriate metrics to evaluate a semantic model via the aforementioned dataset. |
| Outcome: | The proposed protocol is particularly accurate in top-rank evaluation. |
Copied to clipboard
| Challenge: | linguists have discovered patterns which hold across virtually all known natural languages . lingulists are able to learn languages by comparing their learning curves to those of humans . |
| Approach: | They compare LLM learning curves on existing and "impossible" datasets . they find that GPT-2 learns each language and its impossible counterpart equally easily . |
| Outcome: | The proposed model learns each language and its impossible counterpart equally easily, the study shows . the study also shows that the proposed model does not provide any kind of separation between the possible and the impossible . |
Copied to clipboard
| Challenge: | Existing methods for training effective PRMs focus on the first incorrect step and all preceding steps, assuming that all subsequent steps are incorrect. |
| Approach: | They propose a data annotation method specifically designed to score the long CoT reasoning process by using an LLM-based judger for annotation. |
| Outcome: | The proposed method improves PRMs' ability to identify effective self-correction behaviors and reasoning based on erroneous steps. |
Copied to clipboard
| Challenge: | Existing methods for learning knowledge Graphs are incomplete and therefore need well-pretraining. |
| Approach: | They propose a deep reinforcement learning based model which incorporates LSTM and Graph Attention Mechanism as the memory components. |
| Outcome: | The proposed model can get rid of the pretraining process and achieve state-of-the-art performance compared with the other models. |
Copied to clipboard
| Challenge: | Existing metrics for multimodal large language models only focus on token overlap and may not align with human judgment. |
| Approach: | They propose an open-source model that assesses the question answering abilities of multimodal large language models. |
| Outcome: | Experiments show that the ACE-M3 model performs better than existing models and is more reliable than existing metrics. |
Copied to clipboard
| Challenge: | Existing models for knowledge editing focus on knowledge-level or static visual domains, overlooking dynamic semantics. |
| Approach: | They propose a benchmark for modeling large language models using six representative models . they analyze the strengths and limitations of existing models and identify new directions . |
| Outcome: | The proposed benchmark extends existing models from static modalities to dynamic video scenarios. |
Copied to clipboard
| Challenge: | Existing approaches to metric meta-evaluation focus on general statements about absolute and relative quality of metrics across arbitrary system outputs, but in practice, metrics are applied in highly contextual settings. |
| Approach: | They propose a method for contextual metric meta-evaluation by comparing local metric accuracy. |
| Outcome: | The proposed method compares the local metric accuracy of evaluation metrics across translation, speech recognition, and ranking tasks. |
Copied to clipboard
| Challenge: | Pre-trained Transformer models provide robust language representations which can be specialized on various tasks. |
| Approach: | They propose an efficient pruning method based on approximate second-order information that allows pruning weight blocks to be used for pruning. |
| Outcome: | The proposed method is the first to be applied at the BERT scale and significantly pushes the boundaries of the current sparse models with respect to all metrics: model size, inference speed and task accuracy. |
Copied to clipboard
| Challenge: | Neural machine translation models are weak enough for document-level translation . current models only translate sentences individually, resulting in poor document coherence . |
| Approach: | They propose to use the original Transformer model to test document-level neural machine translation . they find that the original transformer models can achieve strong results for document translation if trained properly . |
| Outcome: | The proposed model outperforms sentence-level models on nine datasets and two sentence- level datasets across six languages. |
Copied to clipboard
| Challenge: | Existing evaluations of SAEs focus on metrics such as reconstruction-sparsity tradeoff, human (auto-)interpretability, and feature disentanglement, but they neglect robustness of concept representations to input perturbations. |
| Approach: | They propose an unsupervised approach to map LLM embeddings to sparse interpretable concept embeddables via dictionary learning. |
| Outcome: | The proposed framework shows that sparse autoencoders can manipulate concept-based interpretations without denoising or postprocessing. |
Copied to clipboard
| Challenge: | Existing sequence generation models produce outputs in one pass, usually left-to-right . current models model only a single edit step, and do not fully model editing . |
| Approach: | They propose to model editing processes, modeling the whole process of iteratively generating sequences. |
| Outcome: | The proposed model improves performance on a variety of axes compared to previous models . iterative refinement and editing are central parts of human creative workflow . |
Copied to clipboard
| Challenge: | Grounded text generation systems often generate factual inconsistencies, hindering their real-world applicability. |
| Approach: | They propose a method to assess factual consistency metrics on standardized texts . they recommend NLI and question generation-and-answering-based methods as starting points . |
| Outcome: | The proposed method is more actionable and interpretable than previous methods. |
Copied to clipboard
| Challenge: | Existing methods to evaluate captions have limited learning of their output . previous methods focused on n-gram measures of similarity to reference output based on a ngram of similarities to the output metric. |
| Approach: | They propose a first discourse-aware learned generation metric for evaluating image descriptions. |
| Outcome: | The proposed metric predicts human ratings of captions on out-of-domain images. |
Copied to clipboard
| Challenge: | a recent study shows that noisy reference summaries can be detrimental to model performance. |
| Approach: | They propose to selectively re-write unsupported reference sentences to better reflect source data. |
| Outcome: | The proposed method improves reference quality while retaining all data. |
Copied to clipboard
| Challenge: | a new method to quantify polysemy is based on basic geometry in the contextual embedding space . word sense annotation has always been one of the tasks with the lowest interannotator agreement . |
| Approach: | They propose a method to estimate polysemy based on simple geometry in contextual embedding space. |
| Outcome: | The proposed method is fully unsupervised and data-driven . it can be used to sample sentences with different senses at no extra cost . |
Copied to clipboard
| Challenge: | Existing work in this direction focuses on generating content for standard platforms like Wikipedia, where the content style is fairly consistent, but there could be multiple representations of the same information across the repository. |
| Approach: | They propose an automatic approach to generate an initial version of the author’s intended text based on an input content snippet. |
| Outcome: | The proposed approach improves performance against baselines on several metrics. |
Copied to clipboard
| Challenge: | Creativity measures that distinguish creativity in one domain fail in others, and different metrics disagree on the same data points. |
| Approach: | They examine, analyze, and compare four representative creativity measures across the diverse creative domains, including creative writing, unconventional problem-solving, and research ideation. |
| Outcome: | The measures of creativity across creative domains are compared using a set of human-aligned examples and lack consistency across domains and metrics. |
Copied to clipboard
| Challenge: | Abstractive summarization models have seen great improvements in recent years, but there is limited understanding of the strategies different models employ and how they relate their understanding of language. |
| Approach: | They characterize how one popular abstractive model uses an explicit copy/generation switch to control its level of abstraction vs extraction . they find that abstractive summarization models lack the semantic understanding necessary to generate paraphrases that are both abstractive and faithful to the source document. |
| Outcome: | The proposed model uses syntactic boundaries to truncate sentences that are often copied verbatim. |
Copied to clipboard
| Challenge: | Hallucination remains a critical challenge in large language models (LLMs) in high-stake domains such as legal question answering. |
| Approach: | They propose a method to mitigate hallucination in legal question answering by using behavior cloning and a novel Hard Sample-aware Direct Preference Optimization. |
| Outcome: | The proposed method improves non-hallucinated Statute Rate, Statute Relevance Rate, Legal Claim Truthfulness, and traditional metrics. |
Copied to clipboard
| Challenge: | State-of-the-art summarization systems are trained on massive datasets scraped from the web. |
| Approach: | They manually analyse 600 samples from three popular summarization datasets . they use a six-class typology which captures different noise types and degrees of summarizing difficulty. |
| Outcome: | The proposed model performs better on large datasets than on the current models. |
Copied to clipboard
| Challenge: | Existing referenceless metrics do not take context into account, whereas contextual information is highly valued by BLV users. |
| Approach: | They propose a contextual version of the referenceless metric CLIPScore which addresses the disconnect to the BLV data. |
| Outcome: | The proposed evaluation metrics are based on a proof-of-concept with blind and low vision (BLV) participants. |
Copied to clipboard
| Challenge: | Existing studies have shown that relying on LLMs as information providers may hurt student learning. |
| Approach: | They introduce and apply two bias score metrics to evaluate LLMs for bias in the personalized educational setting, specifically on the models’ roles as “teachers.” |
| Outcome: | The proposed models harm student learning by perpetuating harmful stereotypes and reversing them. |
Copied to clipboard
| Challenge: | a wide consensus is rife regarding the need for reference annotated datasets . however, the creation of such datasets is accompanied by theorectical and practical issues . |
| Approach: | They propose to use agreement among annotators as an indicator of consensus . they argue that it is difficult to produce gold-standard annotated datasets . |
| Outcome: | The proposed model focuses on the complex relations between agreement and reference and the emergence of consensus. |
Copied to clipboard
| Challenge: | Using automated analysis of connected speech is a promising direction for diagnosing cognitive impairments. |
| Approach: | They propose to use a novel model to segment impaired speech transcriptions . they propose to include a Linear Chain CRF and a self-attention mechanism . |
| Outcome: | The proposed system performs better than the existing model with three new datasets used to diagnose cognitive impairments. |
Copied to clipboard
| Challenge: | Neural image-to-text radiology report generation systems have been successful on NLG metrics, but they are not factually complete or consistent due to inadequate training and evaluation. |
| Approach: | They propose a method to improve the factual completeness and correctness of generated radiology reports by using a dataset containing annotated chest X-ray images. |
| Outcome: | The proposed method significantly improves factual completeness and correctness of generated radiology reports on two open radiology report datasets. |
Copied to clipboard
| Challenge: | Existing models of language understanding are based on explicit representations of hierarchical structure, but there are good reasons to doubt that they can be said to understand language in any meaningful way. |
| Approach: | They examine whether syntactic and semantic graph representations can complement and improve neural language modeling. |
| Outcome: | The proposed model outperforms pretrained models on English WSJ in perplexity and other metrics. |
Copied to clipboard
| Challenge: | Existing algorithms to unlearn knowledge and capabilities from large datasets are unclear how to best formulate the unlearning problem. |
| Approach: | They propose to model the hierarchical structure of the unlearning problem, where the forget problem takes priority over the retain problem, and propose an algorithm that aims to unlearn knowledge and capabilities. |
| Outcome: | The proposed algorithm outperforms all state-of-the-art algorithms across unlearning tasks, models, and metrics. |
Copied to clipboard
| Challenge: | Existing evaluation metrics that are not robust to dialect variation are difficult to measure for many groups of users and can penalize systems for producing text in lower-resource dialects. |
| Approach: | They propose a dialect-robust evaluation metric that produces the same score for system outputs that share the same semantics but are expressed in different dialects. |
| Outcome: | The proposed method significantly improves dialect robustness while preserving the correlation between automated metrics and human ratings. |
Copied to clipboard
| Challenge: | Existing approaches to extractive summarization use transformers to learn the structure of long inputs. |
| Approach: | They propose encoder-centric stepwise models for extractive summarization using structured transformers – HiBERT and Extended Transformers . |
| Outcome: | The proposed models outperform previous models on CNN/DailyMail extractive summarization and Rotowire table-to-text generation. |
Copied to clipboard
| Challenge: | Hallucination is a critical challenge for large language models and large vision-language models (LVLMs) however, dedicated research on medical hallucinations remains unexplored. |
| Approach: | They provide a unified perspective on medical hallucination for both LLMs and LVLMs, and delve into its causes. |
| Outcome: | The proposed models have demonstrated impressive performance on a variety of medical benchmarks. |
Copied to clipboard
| Challenge: | Existing evaluation methods for Open Domain Event Detection (ODED) lack representative representations of the real world, making it difficult to accurately reflect performance of various ODED methods in real-world scenarios. |
| Approach: | They propose a scalable and reliable Semantic-level Evaluation framework for Open domain event detection by constructing a more representative evaluation benchmark and introducing a semantic evaluation metric. |
| Outcome: | The proposed framework first constructs a more representative evaluation benchmark that currently includes 564 event types covering 7 major domains, with a cost-effective supplementary annotation strategy to ensure the benchmark’s representativeness. |
Copied to clipboard
| Challenge: | Pre-trained multilingual language models represent multiple languages in a single vector space, a feature hypothesized to enable impressive crosslingual transfer capabilities. |
| Approach: | They propose to use a multilingual representation space that sorts axes based on their language-separability to determine whether geometric distances between languages correlate with crosslingual transfer performance. |
| Outcome: | The proposed measures do not generalize well across models, layers, and tasks. |
Copied to clipboard
| Challenge: | Existing methods to mitigate hallucinations in siMT generate fluency but unfaithful translation. |
| Approach: | They propose a method that utilizes the OMT model to mitigate hallucinations in SiMT. |
| Outcome: | The proposed method reduces hallucinations and improves the SiMT performance. |
Copied to clipboard
| Challenge: | Existing evaluation metrics are not designed to cope with this flexibility. |
| Approach: | They propose to group the qualities into three groups to obtain a single metric called USL-H. |
| Outcome: | The proposed metric achieves good correlations with human judgment and maintains its configurability towards different aspects and metrics. |
Copied to clipboard
| Challenge: | Creating an appealing heading is crucial for attracting readers and marketing work or products. |
| Approach: | They propose a benchmark to measure the quality of heading generation using summarization, neology, and algorithm metrics. |
| Outcome: | The proposed benchmark compared 6,653 abstracts with corresponding descriptions and acronyms and found that it excels across summarization, neology, and algorithm aspects. |
Copied to clipboard
| Challenge: | Existing methods for large language models struggle when the average precision drops below four bits, limiting deployment on resourceconstrained devices such as mobiles, edge sensors, or standard GPUs. |
| Approach: | They propose a game-like game-inspired mixed-precision quantization method which translates these Shapley estimates into a binary quadratic optimization formulation, assigning either 2 or 4-bit precision to layers under strict memory constraints. |
| Outcome: | The proposed method reduces Perplexity by 20 – 80 % across average precisions spanning 4 bit down to 2 bit, compared to methods relying on isolated metrics. |
Copied to clipboard
| Challenge: | Existing metrics for Simultaneous speech translation (SimulST) are inaccurately measuring latency in unsegmented streaming settings. |
| Approach: | They propose to modify existing metrics to correctly measure computation-aware latency for SimulST systems, addressing limitations present in existing metrics. |
| Outcome: | The proposed model is based on a real-time, lowlatency scenario where the model starts generating the textual translation before the entire audio input is processed. |
Copied to clipboard
| Challenge: | Existing methods to incorporate information from other modality, usually static images, are not considered relative to multimodal machine translation. |
| Approach: | They propose a multimodal self-attention method which learns the representation of images based on the text, which avoids encoding irrelevant information in images. |
| Outcome: | The proposed model outperforms previous studies and competitive baselines in terms of various metrics. |
Copied to clipboard
| Challenge: | Currently, standard methods for style transfer have several significant problems. |
| Approach: | They propose to take BLEU between input and human-written reformulations into consideration for benchmarks. |
| Outcome: | The proposed architectures outperform state-of-the-art in style transfer metric on human-written reformulations and take BLEU between input and output into consideration for benchmarks. |
Copied to clipboard
| Challenge: | Data cleaning is a time-consuming and error-prone manual process even with modern workflow tools like OpenRefine. |
| Approach: | AutoDCWorkflow generates a table with a data analysis purpose and generates an open-refine workflow. |
| Outcome: | The proposed pipeline generates clean, minimal tables for data analysis tasks. |
Copied to clipboard
| Challenge: | generating code from a natural language description is a pressing and significant challenge in code intelligence. |
| Approach: | They propose to survey 27 existing large language models for NL2Code and compare them to humanEval benchmarks. |
| Outcome: | The proposed model is compared with existing models on the HumanEval benchmark. |
Copied to clipboard
| Challenge: | Existing approaches to long-term dialogue memory management fail to capture the natural semantic structure of conversations, leading to fragmented and incomplete representations. |
| Approach: | They propose a mechanism that integrates forward- and backward-looking reflections into a personalized memory bank for effective future retrieval. |
| Outcome: | The proposed mechanism outperforms state-of-the-art benchmarks on a long-term dialogue memory model. |
Copied to clipboard
| Challenge: | a number of studies have questioned assumptions of majority vote aggregated labels. |
| Approach: | They construct a model that predicts individual annotator ratings on potentially offensive text and combines this information with the predicted target group of the text to predict the ratings of target group members. |
| Outcome: | The proposed model raises performance over baseline by 22% and 33% at predicting variance among annotators. |
Copied to clipboard
| Challenge: | Current models for dialogue summarization have flaws that may not be well exposed by frequently used metrics such as ROUGE. |
| Approach: | They propose to re-evaluate 18 categories of metrics in terms of four dimensions: coherence, consistency, fluency and relevance, as well as a unified human evaluation of various models for the first time. |
| Outcome: | The proposed dataset will be used to evaluate 18 categories of metrics in terms of coherence, consistency, fluency and relevance, and a unified human evaluation of various models for the first time. |
Copied to clipboard
| Challenge: | MT and diacritization influence performance in a multi-task learning setting, but keeping diacritics is harmful for some languages. |
| Approach: | They propose two classes of metrics to measure the complexity of a diacritical system and propose to use them to compare performance. |
| Outcome: | The proposed metrics correlate positively with the performance of the models. |
Copied to clipboard
| Challenge: | Multimodal sentiment analysis aims to predict the sentiment of video content. |
| Approach: | They propose a framework that performs contrastive representation learning and contrastive feature decomposition to enhance the representation of multimodal information. |
| Outcome: | The proposed framework outperforms baseline methods on CH-SIMS, MOSI and MOSEI datasets on a range of metrics. |
Copied to clipboard
| Challenge: | Recent studies highlight the effectiveness of using in-context learning (ICL) to steer large language models in processing tabular data. |
| Approach: | They propose a method that uses clustering and evolutionary strategies to curate a representative sample set from training data. |
| Outcome: | The proposed method significantly improves fairness across various metrics, showing its efficacy in real-world scenarios. |
Copied to clipboard
| Challenge: | Large language models are prone to contextual hallucination, generating information that is either unsubstantiated or contradictory to the given context. |
| Approach: | They propose a dataset specifically designed for long-context hallucination detection. |
| Outcome: | The proposed architecture outperforms existing models while providing faster inference. |
Copied to clipboard
| Challenge: | Existing image captioning metrics do not capture image relevance . current metrics only measure similarity to ground truth captions . |
| Approach: | They propose a new image relevance metric to evaluate captioning models with veridical visual labels and assess their rate of object hallucination. |
| Outcome: | The proposed metrics show that models with veridical visual labels have higher hallucination rates than models with lower hallucinosity. |
Copied to clipboard
| Challenge: | Prior research focused on identifying the best-performing method to varying hyperparameters . prior research focused primarily on a grid search, which can be impractical for general practitioners . |
| Approach: | They propose a preference optimization method that is more stable across hyperparameters and reduces the average response length. |
| Outcome: | The proposed method increases likelihood of achieving better results through various metrics, such as KL divergence and response length. |
Copied to clipboard
| Challenge: | Retrieval-Augmented Generation allows to enhance Large Language Models with external knowledge. |
| Approach: | They propose a library that allows to benchmark and standardize RAG experiments. |
| Outcome: | The proposed library is an end-to-end library for reproducible research standardizing RAG experiments. |
Copied to clipboard
| Challenge: | Existing work in this field has looked most commonly into gender bias, racial bias, and religious bias. |
| Approach: | They propose an algorithm that uses a neural network to perform ‘soft debiasing’ and build on the seminal work of (CITATION) and (CitATION). |
| Outcome: | The proposed algorithm outperforms current methods on gender, race, and religion metrics on a wide range of metrics. |
Copied to clipboard
| Challenge: | Existing anomaly detection methods require previous observations to be effective . contaminated observations are often not observed, making them ineffective . |
| Approach: | They propose a method that adapts a zero-shot anomaly detector to contaminated observations . they propose an evaluation suite consisting of evaluation protocols and metrics . |
| Outcome: | The proposed method adapts the zero-shot anomaly detector to contaminated observations. |
Copied to clipboard
| Challenge: | Experimental results show that our method not only has a good generalization but also outperforms previous methods on several metrics: BLEU, Content Selection, Content Ordering. |
| Approach: | They propose to build an entity graph from the input tables and introduce a reasoning module to perform reasoning on the graph. |
| Outcome: | The proposed method outperforms previous methods on several metrics: BLEU, Content Selection, Content Ordering. |
Copied to clipboard
| Challenge: | evaluating the clinical quality of medical domain automated text generation remains a challenge. |
| Approach: | They propose a framework for histopathology automated report evaluation that prioritizes clinically relevant content by aligning critical histo pathology entities and relations between reference and generated reports. |
| Outcome: | The proposed framework outperforms existing metrics in histopathology report evaluations. |
Copied to clipboard
| Challenge: | a topic model that incorporates structural relationships connecting documents in socially generated corpora is of limited application in the sciences. |
| Approach: | They propose a topic model that incorporates structural relationships connecting documents in socially generated corpora, such as online forums. |
| Outcome: | The proposed model captures discursive interactions along observed reply links and integrates latent distributed representations in a deep architecture. |
Copied to clipboard
| Challenge: | Existing automatic metrics are observed to correlate poorly with human evaluation. |
| Approach: | They propose to use OpenMEVA to evaluate open-ended story generation metrics. |
| Outcome: | The proposed test suite assesses the capabilities of open-ended story generation metrics on annotated stories and auto-constructed test examples. |
Copied to clipboard
| Challenge: | Existing evaluation metrics are compared based on their ability to correlate with humans, but they disagree in the higher-scoring range in which current systems operate. |
| Approach: | They show that evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate. |
| Outcome: | The evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate. |
Copied to clipboard
| Challenge: | Sexism is a form of oppression based on one's sex and is reported online in numerous ways. |
| Approach: | They propose a multi-task approach for fine-grained multi-label sexism classification that leverages several supporting tasks without incurring manual labeling cost. |
| Outcome: | The proposed method outperforms the state-of-the-art for multi-label sexism classification on a recently released dataset across five standard metrics. |
Copied to clipboard
| Challenge: | Existing theories of language learning for infants are inadequate, according to Chomsky . infants learn language in impoverished environments, according a new study . |
| Approach: | They designed a series of tasks, scenarios, and metrics to simulate the POS . they found that the emerging speech model wav2vec2.0 can learn well in noisy Mandarin environments. |
| Outcome: | The proposed model can learn in noisy and sparse Mandarin environments. |
Copied to clipboard
| Challenge: | Recent work on grouping together views about tweets expressing opinions about the same entities has been criticized for their lack of thematic coherence. |
| Approach: | They propose to use a corpus of microblogs representing opinions about the same topics within the same time window to evaluate thematic coherence. |
| Outcome: | The proposed method outperforms surface level metrics, topic model coherence and text generation metrics (TGMs) but is not as reliable as TGMs due to being less sensitive to time windows. |
Copied to clipboard
| Challenge: | Scientific paper summarization is the focus of recent research . prevailing summarizing methods involve selective extraction of content from abstract, introduction, and conclusion segments within the target articles. |
| Approach: | They propose a model that incorporates references and citations to capture the impact of the document on the research community. |
| Outcome: | The proposed model generates extractive and abstractive summaries in parallel and improves their performance when considering the standard metrics. |
Copied to clipboard
| Challenge: | Large language models are increasingly used to automate data analysis, but data science tasks often admit multiple statistically valid solutions. |
| Approach: | They propose a framework to evaluate LLM-generated code and assess its reproducibility . they introduce two reproducibility-enhancing prompting strategies and benchmark them against standard prompting . |
| Outcome: | The proposed framework improves reproducibility of large language models . it provides a foundation for transparent, reliable, and efficient human–AI collaboration in data science. |
Copied to clipboard
| Challenge: | Existing methods for quantizing large language models focus on breaking down the problem into layer-wise sub-problems and minimizing per-layer error, but this approach lacks theoretical justification and the metrics employed may be sub-optimal. |
| Approach: | They propose a "linearity theorem" establishing a direct relationship between the layer-wise reconstruction error and the model perplexity increase due to quantization. |
| Outcome: | The proposed method outperforms previous data-free methods and improves accuracy-compression trade-offs on Llama-family models. |
Copied to clipboard
| Challenge: | Multimodal speech synthesis is a key challenge due to the scarcity of datasets that pair audio with corresponding video. |
| Approach: | They propose a method that incorporates modality alignment during the pre-training phase on multimodal datasets and freezes the video modality extraction component and the encoder module within the pretrained weights. |
| Outcome: | The proposed method achieves a reduced word error rate (WER) of 31.73%, surpassing the previous best of 33.9% with single-modality audio. |
Copied to clipboard
| Challenge: | Language models are increasingly being studied as models of human language learners. |
| Approach: | They propose a distributional approach to word learning that captures distributional knowledge and gradient preferences for the word’s appropriateness. |
| Outcome: | The proposed signatures capture knowledge of where the target word can and cannot occur as well as gradient preferences about the word’s appropriateness. |
Copied to clipboard
| Challenge: | Document-grounded dialogue systems aim to answer user queries by leveraging external information. |
| Approach: | They propose a dataset to evaluate QA systems' ability to interpret and use structured lists . they use language models and model-based filtering processes to enhance data quality . |
| Outcome: | The proposed model outperforms baselines on the LIST2QA dataset . it shows that the proposed model is more accurate and complete than baselines . |
Copied to clipboard
| Challenge: | Large language models (LLMs) are prone to inconsistencies and individual biases, limiting their reliability. |
| Approach: | They propose a framework that combines ensemble methods with code refinement methodology to address these challenges. |
| Outcome: | The proposed framework outperforms large language models and LLMs with a low-rank averaging and a moderator-based mechanism to simulate human consensus. |
Copied to clipboard
| Challenge: | Language models trained with Maximum Likelihood Estimation (MLE) have been considered as a mainstream solution in Natural Language Generation (NLG) however, they are reportedly suffering from training instability and mode collapse, and therefore outperform conventional MLE models. |
| Approach: | They propose a method to improve Generative Adversarial Nets (GANs) using best student forcing and discriminators to increase training stability and sample diversity. |
| Outcome: | The proposed techniques outperform MLE models and outperformed existing approaches in terms of sample diversity and training stability. |
Copied to clipboard
| Challenge: | Text style transfer is a challenging research task which modifies the linguistic style of a text to meet pre-set objectives such as making the text simpler or more accessible. |
| Approach: | They propose to use large language models to rewrite Dutch news tweets to match specific linguistic styles to achieve a more accessible and accessible text. |
| Outcome: | The proposed prompting strategies perform best for rewriting Dutch news tweets in specific linguistic styles (formal, casual and factual). |
Copied to clipboard
| Challenge: | Hate speech cannot be identified based solely on the presence of specific words; model should reason like humans and be explainable. |
| Approach: | They propose to use Masked Rationale Prediction to predict masked human rationales . the method performs hate speech detection robustly in terms of bias and explainability . |
| Outcome: | The proposed method performs state-of-the-art in terms of bias and explainability. |
Copied to clipboard
| Challenge: | interacting with a model for Visual Question Answering (VQA) quickly reveals that these models lack consistency. |
| Approach: | They propose a dataset, ConVQA, and metrics that enable quantitative evaluation of consistency in VQA. |
| Outcome: | The proposed data augmentation module improves the consistency of VQA models on the Con-VQA dataset and is a strong baseline for future research. |
Copied to clipboard
| Challenge: | Existing models of robustness evaluation are incomprehensive, impractical, and invalid . |
| Approach: | They propose a framework for automatic robustness evaluation that shifts towards model-centric evaluation to further exploit the advantages of adversarial attacks. |
| Outcome: | The proposed framework is based on a model-centric evaluation protocol and a robustness evaluation protocol. |
Copied to clipboard
| Challenge: | GitHub Copilot generates 46% of the code on GitHub. |
| Approach: | They propose a reference-free metric that uses Contrastive Learning to generate meaningful embeddings for code and natural language task descriptions. |
| Outcome: | This paper compares the performance of a new similarity score with existing metrics. |
Copied to clipboard
| Challenge: | Paraphrase generation is an interesting and challenging task which has numerous practical applications. |
| Approach: | They analyze datasets commonly used for paraphrase generation research and show that simply parroting input sentences surpasses state-of-the-art models when evaluated on standard metrics. |
| Outcome: | The proposed model can generate paraphrases even without making any changes to the input sentence or even none at all, compared with other models. |
Copied to clipboard
| Challenge: | Existing methods focus on whether the reasoning chain leads to the correct conclusion, but this view may confound reasoning quality with other spurious shortcuts to predict the answer. |
| Approach: | They propose a framework that evaluates reasoning chains via two key properties: (1) correctness, i.e., each step makes a valid inference based on information contained within the step, preceding steps, and input context, and (2) informativeness, respectively. |
| Outcome: | The proposed framework evaluates reasoning chains via two key properties: (1) correctness, i.e., each step makes a valid inference based on information contained within the step, preceding steps, and input context, and (2) informativeness, which is helpful towards deriving the generated answer. |
Copied to clipboard
| Challenge: | Embodied dialogue instruction following requires an agent to complete a complex sequence of tasks from a natural language exchange. |
| Approach: | They argue that imitation learning and low-level metrics are misleading . they compare existing models with IL and argue evaluation should focus on higher-level semantic goals . |
| Outcome: | The proposed model evaluations are based on three models and compare them with benchmarks . they show that existing models fail to ground query utterances, which are essential for task completion . |
Copied to clipboard
| Challenge: | Recent neural approaches to event temporal relation extraction map events to embeddings in the Euclidean space and train a classifier to detect temporal relations between event pairs. |
| Approach: | They propose to embed events into hyperbolic spaces to model hierarchical structures . they propose to use hyperbolical embeddings to directly infer event relations . |
| Outcome: | The proposed architecture is based on two approaches to encode events and their temporal relations in hyperbolic spaces. |
Copied to clipboard
| Challenge: | Existing methods on diagram generation with LLMs rely heavily on proprietary LLM systems. |
| Approach: | They propose a new evaluation metric to assess demonstration diagrams generated by large language models. |
| Outcome: | The proposed evaluation metric evaluates diagrams produced by state-of-the-art LLMs on recent research literature. |
Copied to clipboard
| Challenge: | Abstractive summarization systems still include factual errors in generated summaries despite recent improvements in factuality detection . |
| Approach: | They aggregate factuality error annotations from nine existing datasets and stratify them according to the underlying summarization model. |
| Outcome: | The proposed method improves on the ChatGPT-based model and shows that it is not superior for all error types. |
Copied to clipboard
| Challenge: | Open Information Extraction (Open IE) systems have been evaluated traditionally via manual annotation. |
| Approach: | They propose to use a dataset to score Open IE systems by matching system predictions with benchmark datasets. |
| Outcome: | The proposed framework matches predictions with the benchmark dataset and is noisy and inconsistent. |
Copied to clipboard
| Challenge: | a goal-driven collaborative drawing task combines language, perception, and actions in a partially observable environment . et al., 1990: 138K messages exchanged between human players. |
| Approach: | They propose a goal-driven collaborative task that combines language, perception, and action . they collect a clip art dataset and use it to build an image-drawing game between two agents . |
| Outcome: | The proposed task integrates language, perception, and action in a virtual world . it is based on a dataset of 10K dialogs and 138K messages exchanged between humans . |
Copied to clipboard
| Challenge: | a new study examines whether beam search can be replaced by a more powerful metric-driven search technique. |
| Approach: | They propose a beam search method which is agnostic to the end metric and report results on a variety of metrics. |
| Outcome: | The proposed method is based on a Monte-Carlo Tree Search (MCTS) based method and shows it can be used in language applications. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) must possess an understanding of the nation’s culture and basic knowledge. |
| Approach: | They propose to construct a national alignment benchmark, KorNAT, which measures the alignment between an LLM and a targeted country from two perspectives: social value alignment and common knowledge alignment. |
| Outcome: | The proposed model passes the national alignment score of 7 LLMs, indicating there is room for improvement. |
Copied to clipboard
| Challenge: | Existing summarization models produce unfaithful outputs for medical text summarizing . a framework to improve faithfulness is proposed to improve medical text summary accuracy . |
| Approach: | They propose a framework to improve faithfulness by fine-tuning pre-trained language models based on medical knowledge. |
| Outcome: | The proposed framework improves faithfulness on medical summarization tasks. |
Copied to clipboard
| Challenge: | Existing methods for text generation evaluation metrics are lacking in robustness analysis. |
| Approach: | They propose to use stress tests to test for errors in text generation evaluation metrics . they find that BERTScore is confused by truncation errors in summarization . |
| Outcome: | The proposed stress tests show that they are insensitive to errors in open-ended generation, translation, and summarization. |
Copied to clipboard
| Challenge: | Existing work suggests that the degree of hallucination depends on factual errors in training data. |
| Approach: | They propose a method to use training data to reduce hallucination by ensembling parameter variations in training data. |
| Outcome: | The proposed method improves on XSUM and CNN/DM datasets on human evaluations and factual metrics. |
Copied to clipboard
| Challenge: | a new paper aims to reproduce the work described in Vajjala & Rama (2018) . the paper focuses on features-based and neural approaches to essay scoring in Czech, German and Italian . |
| Approach: | They propose to replicate the work described in Vajjala & Rama 2018, ‘Experiments with universal CEFR classification’, as part of REPROLANG 2020. |
| Outcome: | The proposed methods perform better than feature-based models for large text datasets, though neural network modifications do bring performance closer to the best feature-driven models. |
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) training is based on document-level metrics, not sentence-level BLEU. |
| Approach: | They propose to merge document-level metrics with batch-level documents to improve NMT training. |
| Outcome: | The proposed training is more robust for document-level metrics than sequence MRT and maximum-likelihood training. |
Copied to clipboard
| Challenge: | Existing models for multi-hop reasoning are not able to evaluate their interpretability . a recent study found that many paths are unreasonable . |
| Approach: | They propose a framework to evaluate the interpretability of multi-hop reasoning models . they annotate all possible rules and establish a benchmark . |
| Outcome: | The proposed framework outperforms existing models in terms of performance and interpretability. |
Copied to clipboard
| Challenge: | Similes are a crucial part of creative writing, but there is still a lack of evaluation metrics for simile generation. |
| Approach: | They propose to use similes as a tool to evaluate simile generation metrics . they propose to combine five criteria and automatic metrics for each criterion . |
| Outcome: | The proposed metrics are significantly more correlated with human ratings from each perspective compared with prior automatic metrics. |
Copied to clipboard
| Challenge: | Existing evaluation frameworks for text summarization lack domain-specific assessment criteria and are predominantly English-centric. |
| Approach: | They propose a multi-dimensional, multi-domain evaluation of summarization in English and Chinese that incorporates specialized assessment criteria for each domain and leverages a debate system to enhance annotation quality. |
| Outcome: | The proposed evaluation framework provides a multi-dimensional, multi-domain evaluation of summarization in English and Chinese. |
Copied to clipboard
| Challenge: | Existing benchmarks measure the correlation with human judgements of faithfulness on model-generated summaries, but they are insufficient for diagnosing whether metrics are consistent, effective on human-written texts, and sensitive to different error types. |
| Approach: | They propose to use unfaithful minimal pairs to measure the consistency of automatic faithfulness metrics by comparing human-written summary pairs with a dataset of 889 human-writing, minimally different summary pairs. |
| Outcome: | The proposed benchmarks show that the most discriminative metrics tend not to be the most consistent, and that the best performing metrics are sensitive to errors. |
Copied to clipboard
| Challenge: | a gap between conversations can be weeks, months or years, and dialogue systems which do not explicitly model time may generate unnatural responses. |
| Approach: | They propose to model the passage of time between conversations by exposing time information to a multi-session dialogue dataset and comparing different representations of time and event progress. |
| Outcome: | The proposed model is based on a real-time dataset showing that it can predict topics and information gained from conversations over a long time span. |
Copied to clipboard
| Challenge: | Existing speaker models learn strategies to evade evaluation metrics and obtain higher scores even for low-quality sentences. |
| Approach: | They propose a speaker-based instruction generator that utilises both structural and semantic knowledge of the environment to produce richer instructions. |
| Outcome: | The proposed model outperforms existing models and is evaluated using standard metrics. |
Copied to clipboard
| Challenge: | Existing metrics for open-ended text generation are poorly correlated with human judgments . despite the success of existing metrics, there are few plausible outputs for the same input . |
| Approach: | They propose a UNreferenced measure for evaluating open-ended story generation . it is built on top of BERT and is trained to distinguish human-written stories from negative samples . |
| Outcome: | The proposed measure is more generalizable than state-of-the-art metrics and correlates better with human judgments. |
Copied to clipboard
| Challenge: | ECTSum is a dataset for bullet-point summarization of earnings calls hosted by publicly traded companies. |
| Approach: | They propose a dataset with transcripts of earnings calls and bullet point summaries derived from Reuters articles. |
| Outcome: | The proposed dataset compares transcripts of earnings calls hosted by publicly traded companies with experts-written bullet point summaries derived from Reuters articles . |
Copied to clipboard
| Challenge: | Existing methods to learn compact cluster representations from coarsely labeled data are noisy and degrade the quality of learning. |
| Approach: | They propose a framework that encodes semantic structures of data into the embedding space . they retrieve k-nearest neighbors of a query as positive keys to capture similarities . |
| Outcome: | The proposed framework can retrieve more accurate neighbors and outperform state-of-the-art models by a large margin. |
Copied to clipboard
| Challenge: | a systematic way to measure agreement across bias metrics and models is lacking . a lack of agreement between metrics and model results may be a problem . |
| Approach: | They introduce Metric Agreement Score and Model Agreement Score to measure agreement across bias metrics and models. |
| Outcome: | The proposed measures show that metrics within the same category behave independently of each other. |
Copied to clipboard
| Challenge: | Traditionally, turn-taking is done using a simple silence threshold, but more modern approaches use cues known to be important in human-human turn-shifts. |
| Approach: | They propose a turn-taking and response-ranking model that conditions the end-of-turn prediction on conversation history and what the next speaker wants to say. |
| Outcome: | The proposed model outperforms the baseline model in a variety of metrics. |
Copied to clipboard
| Challenge: | a recent study shows that open-source large language models (LLMs) exhibit diverse strengths and weaknesses due to variations in their architectures and training data. |
| Approach: | They propose a framework that leverages the diverse strengths of open-source large language models. |
| Outcome: | The proposed framework outperforms individual LLMs and baseline methods across various metrics, establishing a substantial performance gap. |
Copied to clipboard
| Challenge: | a recent human evaluation study found that translations produced by current MT systems achieve very high-quality scores when judged by humans on a direct assessment scale of 0 to 100. |
| Approach: | They stress-test the ability of current translation quality metrics to detect correct translations . they show that current metrics often over or underestimate translation quality . |
| Outcome: | The proposed method overestimates translation quality, the authors show . they show that current metrics often overestimate translation quality . |
Copied to clipboard
| Challenge: | Existing methods focus on local optimal while ignoring sole-mention disambiguation boosted by richer context from other mentions’ disambiguating processes. |
| Approach: | They propose an approach to extracting medical entity disambiguation using memory mechanism and memorized entity information (M3E) they use a memory mechanism module that performs memory caching, retrieval, fusion and cross-network residual to aid the disambiguations of remaining mentions. |
| Outcome: | The proposed method outperforms state-of-the-art methods on two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods focus on constructing multi-perspective prompts to expand instructions, overlooking the “Fixed Thinking Pattern” issue of Large Language Models. |
| Approach: | They propose a method that analyzes the statistical characteristics of newly generated instructions and updates the prompts after a fixed number of instruction expansions. |
| Outcome: | The proposed method surpasses open-source LLMs and GPT3.5 in several metrics. |
Copied to clipboard
| Challenge: | Existing language models to generate implicit hate explanations are lacking in many fields. |
| Approach: | They propose to use language models to generate explicit hate posts to make it clear . they find that simpler models incorporating external toxicity signals outperform KG-infused models . |
| Outcome: | The proposed setup produces more precise explanations than zero-shot GPT-3.5, highlighting the intricate nature of the task. |
Copied to clipboard
| Challenge: | specialised language models (LMs) have shown to exhibit lower perplexity and higher downstream performance across the board. |
| Approach: | They propose a benchmark for NLP evaluation in social media, SuperTweetEval. |
| Outcome: | The proposed benchmark shows that social media models perform better when compared to general-purpose models, metrics and benchmarks. |
Copied to clipboard
| Challenge: | Current automated RAG metrics perform poorly in clinical and conversational use cases. |
| Approach: | They propose an automated and scaleable TRIaD for evaluating clinical QA systems leveraging Retrieval Augmented Generation (RAG) metric consisting of three metrics: Context Relevance (CR), Refusal Accuracy (RA), and Conversational Faithfulness (CF). |
| Outcome: | The proposed metric captures the faithfulness of a model’s response without penalising conversational elements and captures refusal to address questions outside of the system’s scope of practice. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) suffer from hallucinations and outdated knowledge due to their reliance on static training data. |
| Approach: | They review training strategies, robustness enhancements, loss functions, and agent-based approaches and outline open challenges and future directions to guide research in this evolving field. |
| Outcome: | The proposed model improves accuracy and accuracy while integrating external dynamic information for improved factual grounding. |
Copied to clipboard
| Challenge: | Existing studies rely on single-task evaluations and classification-based metrics that overlook the fundamental differences between generative LLMs and traditional classification models. |
| Approach: | They propose to use four new metrics to evaluate LLM-based word sense disambiguation (WSD) . experimental results reveal significant limitations in LLMs' WSD performance . |
| Outcome: | The proposed evaluation framework is open-source at https://github.com/DayDream405/RoDEval. |
Copied to clipboard
| Challenge: | Large language models excel at factual recall, arithmetic reasoning, multi-turn dialogue . their capacity as askers, formulating strategic, adaptive, and information-seeking questions, remains less explored . |
| Approach: | They propose a protocol for evaluating large language models as strategic question-askers . they propose entropy-based methods that filter candidates via ConceptNet and Bayesian method that tracks belief updates over semantic concepts . |
| Outcome: | The proposed method is model-agnostic and supports post hoc analysis. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are widely used for various tasks but evaluating the consistency of generated text remains a challenge. |
| Approach: | They propose a prompt-based consistency metric which provides explanations for consistency scores by providing detailed reasoning and pinpointing inconsistent text spans. |
| Outcome: | The proposed metric outperforms state-of-the-art metrics in summarization, free text generation and data-to-text conversion tasks by 8.7% and 6.2%. |
Copied to clipboard
| Challenge: | Existing evaluation protocols and metrics do not capture the full spectrum of LLM capabilities, especially in complex reasoning tasks. |
| Approach: | They propose a new evaluation metric that continuously assesses model performance across multiple sampling attempts, quantifying both the model’s potential capabilities and operational consistency. |
| Outcome: | The proposed evaluation metric measures model performance across multiple sampling attempts and provides comprehensive insights into their potential capabilities and operational consistency. |
Copied to clipboard
| Challenge: | Existing methods for evaluating statement autoformalization are limited . current methods can achieve up to 45.1% accuracy on undergraduate mathematics . |
| Approach: | They propose a new autoformalization metric that correlates strongly with human judgment . they propose two new auto-formalisation benchmarks: ProofNet# and RLM25 . |
| Outcome: | The proposed methods can achieve up to 45.1% accuracy on undergraduate mathematics but struggle with research-level content without proper context. |
Copied to clipboard
| Challenge: | Generative commonsense reasoning requires models to synthesize coherent narratives that satisfy lexical constraints and commonsensical logic. |
| Approach: | They propose a framework that allows for deep semantic diversity rather than surface-level lexical variation. |
| Outcome: | The proposed framework achieves over 10% improvement in overall accuracy on NoRa and SPICE score on CommonGen-Lite. |
Copied to clipboard
| Challenge: | In-context learning (ICL) is a training-free paradigm of fewshot inference that can generalize to novel tasks by conditioning on a few task examples. |
| Approach: | They show that BERTScore-Recall (BSR) selects better examples that demonstrate more of the salient aspects of the test input. |
| Outcome: | The proposed model outperforms methods that leverage task or LLM-specific training on compositional tasks. |
Copied to clipboard
| Challenge: | Current evaluation practices in Simultaneous Speech Translation systems involve segmenting the input audio and its translations, calculating quality and latency metrics for each segment, and averaging the results. |
| Approach: | They propose to use the mean to estimate latency for Simultaneous Speech Translation systems to provide a better understanding of their results. |
| Outcome: | The proposed methods can provide a better understanding of SimulST systems’ latency. |
Copied to clipboard
| Challenge: | Existing evaluation metrics are unreliable for factual consistency tasks, limiting their effectiveness as signals for shaping model behaviour. |
| Approach: | They propose an automated training pipeline that improves factual consistency in summaries by aggregating scores from different weak metrics. |
| Outcome: | The proposed approach improves factual consistency in summaries by aggregating scores from weak metrics. |
Copied to clipboard
| Challenge: | Existing tools for measuring representational harms caused by large language model systems are not useful for practitioners. |
| Approach: | They examine the extent to which public instruments are used to measure representational harms caused by large language model-based systems. |
| Outcome: | The proposed instruments do not meet the needs of practitioners evaluating large language model-based systems. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated impressive performance in many reasoning tasks, but temporal reasoning remains challenging due to its intrinsic complexity. |
| Approach: | They propose a new prompting technique tailored for temporal reasoning, Narrative-of-Thought (NoT), that first converts the events set to a Python class, then prompts a small model to generate a temporal narrative. |
| Outcome: | The proposed technique achieves the highest F1 on Schema-11 evaluation set, while securing an overall F1 of par with GPT-3.5/4. |
Copied to clipboard
| Challenge: | Existing methods for paragraph captioning videos without event ground truths generate one sentence for each event, but without event labels, it is difficult to locate the transitions between events and minimize repetition. |
| Approach: | They propose a module that dynamically groups event information with the help of action information for the entire video and excludes redundant frames within pre-defined clips. |
| Outcome: | The proposed module outperforms the state-of-the-art methods on all metrics. |
Copied to clipboard
| Challenge: | Existing summarization metrics favor shorter or longer summaries, but evaluations of these metrics are flawed. |
| Approach: | They propose a Bayesian normalization technique that effectively diminishes this bias. |
| Outcome: | The proposed method significantly improves the concordance between human annotators and most metrics in terms of summary coherence. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are expensive yet powerful ways to annotate text, and can be inconsistent when compared with experts. |
| Approach: | They propose to combine LLM annotations with a limited number of expensive expert annotations to produce valid estimates. |
| Outcome: | The proposed methods produce consistent estimates under theoretical assumptions, but they are not comparable across finite datasets. |
Copied to clipboard
| Challenge: | Using the Decompose-Then-Verify framework, such as FActScore, can be manipulated by adding obvious or repetitive subclaims to artificially inflate scores. |
| Approach: | They propose a decomposition-based tool called Core to filter subclaims based on their uniqueness and informativeness. |
| Outcome: | The proposed evaluation framework supports easy and modular use of Core and various decomposition strategies. |
Copied to clipboard
| Challenge: | MATCHA is an automatic metric that rewards semantic agreement with a reference and penalizes contradictions. |
| Approach: | They introduce a metric that jointly rewards semantic agreement with a reference and penalizes contradictions. |
| Outcome: | The proposed metric outperforms popular metrics on eight public benchmarks compared with human annotations on question-answering, image caption generation, natural language inference, summarization, and semantic textual similarity tasks. |
Copied to clipboard
| Challenge: | Persona-prompting is a growing strategy to personalize outputs, but its impact on how LLMs represent social groups remains underexplored. |
| Approach: | They investigate whether persona-prompting leads to different levels of linguistic abstraction . they compare 11 persona driven responses to those of a generic AI assistant . |
| Outcome: | The proposed method can be used to personalize outputs, but its impact on how LLMs represent social groups remains underexplored. |
Copied to clipboard
| Challenge: | Existing studies have shown that automated evaluation approaches correlate with human ratings in English, but this is unclear for other languages. |
| Approach: | They construct a small-scale pilot dataset containing article-summary pairs and human ratings in English, Chinese and Indonesian to measure the strength of summaries. |
| Outcome: | The results show that standard metrics are unreliable measures of quality in Chinese and Indonesian. |
Copied to clipboard
| Challenge: | Existing methods for accelerating Large Vision-Language Models lack comprehensive evaluation across diverse backbones, benchmarks, and metrics. |
| Approach: | They propose EffiVLM-BENCH framework for evaluating absolute performance and generalization and loyalty. |
| Outcome: | The proposed framework offers insights into optimal strategies for accelerating LVLMs. |
Copied to clipboard
| Challenge: | Existing methods to generate LLMs with a single ‘best’ prompt are unstable and sub-optimal in practice. |
| Approach: | They propose to decode multiple candidate generations from a prompt bank at inference-time and use Minimum Bayes Risk (MBR) to select a final output. |
| Outcome: | The proposed method improves MBR across a set of conditional generation tasks and models. |
Copied to clipboard
| Challenge: | Existing approaches focus on downstream metrics to select QA pairs, which lack generalization across different datasets. |
| Approach: | They propose a general selection method that uses a large pre-trained language model as a reward model in a Reinforcement Learning framework for the training of the selection agent. |
| Outcome: | The proposed method improves performance on generative and extractive datasets. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are prone to forced generation when confronting ambiguous evidence or complex recursive dependencies. |
| Approach: | They propose a framework that imposes semantic and structural constraints via a financial metric knowledge graph. |
| Outcome: | a neuro-symbolic framework outperforms existing models on financial metric knowledge graphs. |
Copied to clipboard
| Challenge: | Despite significant advances in video-language modeling, hallucinations remain a persistent challenge in video large language models. |
| Approach: | They present a systematic taxonomy that categorizes hallucinations into two core types: dynamic distortion and content fabrication. |
| Outcome: | The proposed taxonomy categorizes hallucinations into two core types: dynamic distortion and content fabrication. |
Copied to clipboard
| Challenge: | Large language models (LLMs) can learn to perform a wide range of tasks, but generating valid molecules using representations like SMILES is challenging in few-shot settings. |
| Approach: | They propose a language framework that converts invalid SMILES to SELFIES and LLMs as post-hoc correctors to ensure that the molecules generated by LLM are 100% valid. |
| Outcome: | The proposed model performs worse with SELFIES than with SMILES and improves on other metrics. |
Copied to clipboard
| Challenge: | Existing approaches transfer the soft prompt to low-source targets by combining all source tasks or a single “high-similar” source task one-time-only. |
| Approach: | They propose a method to group similar source tasks based on two metrics: target similarity and knowledge consistency. |
| Outcome: | The proposed method reduces negative transfer and improves performance on low-source targets. |
Copied to clipboard
| Challenge: | Current data selection paradigms rely on static, externally defined metrics, which fail to adapt to the evolving capabilities of models during training. |
| Approach: | They propose a dynamic sampling framework that aligns training data with the model's intrinsic competence by iterating on real-time feedback. |
| Outcome: | Extensive experiments on eight benchmarks show that SAI-DPO outperforms static baselines at most nearly 6 points, achieving state-of-the-art efficiency with significantly less data. |
Copied to clipboard
| Challenge: | Existing approaches to producing presentation slides rely on fixed templates or executable code . Existing methods rely only on predefined templates and emit executable codes . |
| Approach: | They propose a hierarchical slides generation workflow DeepSlides that organizes slide design tasks without any predefined template or style. |
| Outcome: | The proposed framework outperforms baseline methods on evaluated metrics and achieves superior performance in human preference evaluations. |
Copied to clipboard
| Challenge: | Existing studies on robustness to explicit noise (e.g., document semantics) but overlook implicit noise (spurious features). |
| Approach: | They propose a framework to quantify the robustness of RAGs against spurious features by integrating a data synthesis pipeline and a taxonomy. |
| Outcome: | The proposed framework quantifies the robustness of RALMs against spurious features. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have superior translation performance and long-context capabilities, but evaluation methodologies remain constrained to sentence-level assessment due to dataset limitations and token number restrictions in metrics. |
| Approach: | They propose an evaluation scheme that extends existing automatic metrics to long-document translation by treating documents as continuous text and applying sentence segmentation and alignment methods. |
| Outcome: | The proposed evaluation scheme outperforms existing long-form document evaluation schemes while accounting for under-/over-translations and varied sentence boundaries. |
Copied to clipboard
| Challenge: | Large language models (LLMs) produce incomplete or selectively omit key information . omissions of key information or misrepresentation of conflicting evidence can cause harm . |
| Approach: | They propose a method that decomposes texts into atomic statements and uses natural language inference to identify missing facts and a Q A-based metric that extracts question-answer pairs and compares responses across sources. |
| Outcome: | The proposed evaluation metrics show they perform better than more complex metrics, but at a cost. |
Copied to clipboard
| Challenge: | Qualitative analysis is widely adopted across many social science disciplines. |
| Approach: | They propose a theory-informed computational method for measuring inductive coding results from humans and GAI. |
| Outcome: | The proposed method captures breadth, consensus, unique contribution, and systematic deviation without assuming ground truth. |