Papers with correlation
Copied to clipboard
| Challenge: | Existing image captioning metrics focus on linguistic aspects and do not match human judgements at sentence-level. |
| Approach: | They propose to incorporate lexical and semantic metrics as features to capture adequacy and fluency of captions at different linguistic levels. |
| Outcome: | The proposed framework captures adequacy and fluency of captions at different linguistic levels. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for text simplification focus on only one dimension: fluency, simplicity and meaning preservation. |
| Approach: | They introduce a dataset to assess legal meaning preservation between two legal texts . they also introduce sanity checks for two identical sentences . |
| Outcome: | The proposed metric shows superior correlation with human judgment compared to existing metrics. |
Copied to clipboard
| Challenge: | Using a method to collect references and compare their value with human evaluations, we show that multi-reference BLEU does not improve the correlation for high quality output. |
| Approach: | They propose a method to compare the quality of automated metrics by analyzing references and comparing them with human evaluations. |
| Outcome: | The proposed method improves correlation with all modern evaluation metrics including embedding-based methods. |
Copied to clipboard
| Challenge: | a novel hate speech detection model can be used to detect word- and character-level adversarial attacks . existing adversarials assume that attackers replace the target words with other names to evade detection . |
| Approach: | They propose a robust hate speech detection model that can defend against adversarial attacks . they describe the process of hate speech recognition by a causal graph and a regularized entropy loss function to quantify spurious correlation . |
| Outcome: | The proposed model can defend against word- and character-level adversarial attacks. |
Copied to clipboard
| Challenge: | Existing studies have failed to explore co-attentive multi-modal modeling for visual and text reasoning. |
| Approach: | They propose to use image and multi-modal Transformers to reconstruct fMRI brain activity . they use two popular datasets to study visual and text reasoning . |
| Outcome: | The proposed model outperforms existing models on two popular datasets . the results raise the question whether visual processing is affected implicitly by linguistic processing . |
Copied to clipboard
| Challenge: | Existing studies have relied on out-of-the-box machine translation metrics to evaluate interpretation data, but they do not account for human judgments of interpretation quality. |
| Approach: | They propose to use machine translation metrics to evaluate human interpretations to address potential barriers to disfluency, summarization, paraphrasing and segmentation. |
| Outcome: | The proposed model achieves better correlation with human judgments than state-of-the-art metrics. |
Copied to clipboard
| Challenge: | evaluating machine translation (MT) with cross-lingual information retrieval is relatively time-consuming and subjective. |
| Approach: | They propose a toolkit that evaluates machine translation with a proxy task of cross-lingual information retrieval. |
| Outcome: | The proposed toolkit is based on the "metrics shared task" of WMT2019. |
Copied to clipboard
| Challenge: | Existing methods for efficiently eliciting scalar annotations for dataset construction and system quality estimation by human judgments are not shown. |
| Approach: | They propose a method for efficiently eliciting scalar annotations by human judgments. |
| Outcome: | The proposed method leads to increased correlation with ground truth, suggesting it is an improved mechanism for dataset creation and manual system evaluation. |
Copied to clipboard
| Challenge: | Existing automated evaluation metrics fail to consider factual correctness or are limited in their interpretability. |
| Approach: | They propose a radiology report evaluation metric that leverages natural language understanding of language models to identify and explain clinically significant errors. |
| Outcome: | The proposed method demonstrates higher correlation with expert error counts and higher alignment with expert preferences when compared to previous methods. |
Copied to clipboard
| Challenge: | a new approach to multilingual word embedding is needed to achieve this goal . a multilingual common semantic space is a language-agnostic semantic continuous space . |
| Approach: | They propose a multilingual common semantic space where words from multiple languages are mapped into a shared space so that resources and knowledge can be shared across languages. |
| Outcome: | The proposed approach achieves 14.6% absolute F-score gain over state-of-the-art methods on cross-lingual direct transfer. |
Copied to clipboard
| Challenge: | Existing approaches require dialog datasets to explicitly annotate knowledge base (KB) queries. |
| Approach: | They propose a pipelined approach to predict when to make a KB query and train the dialog agent without explicit annotation. |
| Outcome: | The proposed approach predicts when to make a KB query, then predicts a query at the predicted position and uses the results in subsequent dialog. |
Copied to clipboard
| Challenge: | Existing methods for graded entity salience are subjective but lack consistency. |
| Approach: | They propose a method for graded entity salience that combines subjective judgments and summarization-based methods that define saliency as mention-worthiness in a summary. |
| Outcome: | The proposed approach outperforms existing methods and shows stronger correlation with human summaries and alignments. |
Copied to clipboard
| Challenge: | Current methods for automated fact-checking rely on relying on other evaluation metrics and closed knowledge sources. |
| Approach: | They propose a method which combines evidence evaluation with verdict-level proxy scoring. |
| Outcome: | The proposed method outperforms existing methods in accuracy and robustness against human ratings and adversarial tests. |
Copied to clipboard
| Challenge: | BLEU and METEOR metrics fail to provide information on which linguistic factors impact performance of natural language generation models. |
| Approach: | They propose a framework for error analysis which permits identifying which features of the input affect the models’ results. |
| Outcome: | The proposed framework improves the performance of 174 system runs submitted to the Multilingual SR shared tasks. |
Copied to clipboard
| Challenge: | Existing studies on multi-document summarization focus on collating information that all sources agree upon, but the task of summarizing diverse information remains underexplored. |
| Approach: | They propose a task of summarizing diverse information encountered in multiple news articles encompassing the same event using a dataset curated by a large language model. |
| Outcome: | The proposed task aims to summarize diverse information in multiple news articles encompassing the same event . the proposed task is difficult due to its limited coverage and verbosity biases . |
Copied to clipboard
| Challenge: | Public companies in the US are required to publish annual reports that contain over 25,000 words across all sections and a high percentage of boilerplate content that does not change much year-to-year. |
| Approach: | They propose to model complex, cross-document relationships between financial reports using paired financial reports. |
| Outcome: | The proposed model can predict company risk and correlation from financial reports . the proposed model is able to recognize complex, nuanced relationships with complex signals . |
Copied to clipboard
| Challenge: | Existing approaches require large amounts of expert annotated data, computation, and time for training. |
| Approach: | They propose an unsupervised approach to QE where no training is required . they use a dataset that enables work on both black-box and glass-box approaches . |
| Outcome: | The proposed approach rivals state-of-the-art supervised QE models in terms of correlation with human judgments of quality. |
Copied to clipboard
| Challenge: | Existing methods do not correlate strongly with human annotations. |
| Approach: | They propose a method that measures the probability that a language model will continue the conversation with a fixed set of follow-ups. |
| Outcome: | The proposed method achieves the highest correlation with human evaluations when compared against twelve existing methods. |
Copied to clipboard
| Challenge: | Text style transfer (TST) is a multidimensional task requiring the assessment of style transfer accuracy, content preservation, and naturalness. |
| Approach: | They propose to use text style transfer metrics to evaluate outputs of text editors . they also investigate the potential of large language models as tools for TST evaluation . |
| Outcome: | The proposed methods provide better insights than existing metrics, the authors show . their meta-evaluation through correlation with hu-man judgments shows they are effective . |
Copied to clipboard
| Challenge: | Existing studies on multilingual image captioning have been hampered by a lack of high-quality evaluation datasets. |
| Approach: | They present a dataset of 3600 images annotated with human-generated captions in 36 languages. |
| Outcome: | The proposed dataset shows that it is feasible to build multilingual image captioning models trained on machine-translated data. |
Copied to clipboard
| Challenge: | evaluators of machine translation systems often use text-based metrics to evaluate performance . however, these metrics lack semantic-level information and exhibit poor correlation with human ratings . authors propose a method to reduce inference bias of neural metrics in out-of-distribution data . |
| Approach: | They propose to reduce inference bias by using uncertainty estimation, test-time adaptation, and inference to reduce model uncertainty. |
| Outcome: | The proposed method reduces model uncertainty and improves correlation performance across models. |
Copied to clipboard
| Challenge: | Existing models for dialog evaluation are trained using a single relevant response and multiple random negatives. |
| Approach: | They propose a dataset to test whether model-based dialog evaluation metrics can be used to train models . they propose n-gram based metrics and embedding based ones to be used for model-driven evaluation . |
| Outcome: | The proposed model outperforms existing models on a reddit dataset on relevant responses and adversarial responses. |
Copied to clipboard
| Challenge: | a neural network estimation system for spoken dialogues can be used to estimate the communication style of a user's interaction, but this is rarely implemented in a live system. |
| Approach: | They propose a neural network approach to estimate the communication style of spoken interaction, namely elaborateness and directness. |
| Outcome: | The proposed method can estimate the elaborateness and directness of spoken interaction and improve the results with additional linguistic features. |
Copied to clipboard
| Challenge: | Neural metrics have a high correlation with human judgements but they are hard to eliminate due to their "black box" nature. |
| Approach: | They propose to use minimum bayes risk decoding to explore and quantify weaknesses in COMET models. |
| Outcome: | The proposed model is not sensitive enough to discrepancies in numbers and named entities, and is hard to remove by training on additional synthetic data. |
Copied to clipboard
| Challenge: | Recent work has attempted to enhance vector space representations using information from structured semantic resources. |
| Approach: | They propose a root-mean-square error evaluation metric to evaluate the utility of different lexical resources for retrofitting. |
| Outcome: | The proposed method improves word similarity performance by using root-mean-square error (RMSE) and root-macro-error (RMME) metric. |
Copied to clipboard
| Challenge: | Reliably evaluating Machine Translation (MT) through automated metrics is a long-standing problem. |
| Approach: | They propose to use MT models to generate multiple diverse translations and use them as surrogates to reference translations to obtain a quantification of translation variability. |
| Outcome: | The proposed approach improves correlation with human judgements of quality by 15%. |
Copied to clipboard
| Challenge: | Patent-CR is the first dataset created for the patent claim revision task in English. |
| Approach: | They propose to create a dataset for the patent claim revision task in English that includes both initial patent applications rejected by examiners and the final granted versions. |
| Outcome: | The proposed dataset includes both initial patent applications rejected by examiners and the final granted versions. |
Copied to clipboard
| Challenge: | Existing evaluation methods for summarization systems measure semantic overlap between a system summary and a human reference on word-string level. |
| Approach: | They propose to use distributed representations to evaluate system summary and human reference on word-string level. |
| Outcome: | The proposed representations outperform ROUGE on recent corpora but are less good on test data used in previous studies. |
Copied to clipboard
| Challenge: | Existing methods to evaluate text summarization tasks using ROUGE have been criticized for lack of semantic understanding. |
| Approach: | They propose a semantic-aware metric for extractive summarization task that is semantic-based . they use CNN/DailyMail dataset to study the new metric . |
| Outcome: | The proposed metric is semantic-aware and shows higher correlation with human judgement and yields a large number of disagreements with the original ROUGE metric. |
Copied to clipboard
| Challenge: | Existing MLLMs are optimized for single-task scenarios and struggle to generalize to diverse contexts. |
| Approach: | They propose a framework that integrates multitask reinforcement learning and generalization capabilities of MLLMs to optimize the judge model across multiple tasks. |
| Outcome: | The proposed framework outperforms baseline models in judgment consistency and correlation with human preferences. |
Copied to clipboard
| Challenge: | Existing evaluation frameworks for natural language generation are dominated by similarity-based metrics. |
| Approach: | They propose a multi-dimensional evaluator for natural language generation that integrates multiple dimensions into one evaluer. |
| Outcome: | The proposed evaluator improves on three typical NLG tasks and improves with external knowledge. |
Copied to clipboard
| Challenge: | Understanding natural language requires common sense, one aspect of which is the ability to discern the plausibility of events. |
| Approach: | They propose a method of forcing model consistency that improves correlation with human plausibility judgements. |
| Outcome: | The proposed method improves correlation with human plausibility judgements. |
Copied to clipboard
| Challenge: | Existing n-gram similarity metrics fail to discriminate the incorrect answers due to the free-form of the answer. |
| Approach: | They propose a new metric that assigns different weights to each token via keyphrase prediction to judge the correctness of GenQA. |
| Outcome: | The proposed metric has a significantly higher correlation with human judgments than existing metrics in various datasets. |
Copied to clipboard
| Challenge: | Visual captioning is an open-ended area for evaluation, requiring specialized training to improve human-correlation. |
| Approach: | They propose a new evaluation framework rooted in information theory . they propose metric SPURTS and metric SMURF to measure fluency . |
| Outcome: | The proposed metrics achieve state-of-the-art correlation with human judgment compared with other evaluation metrics. |
Copied to clipboard
| Challenge: | Existing image captioning metrics provide a single score to measure caption qualities, which are less explainable and informative. |
| Approach: | They propose an Informative Metric for Reference-free Image Caption evaluation to support this feedback . they propose to provide a text precision score, a vision recall score and an overall quality score . |
| Outcome: | The proposed method improves on existing metrics on multiple benchmarks and compares coarse-grained scores with human judgements. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for machine translation are difficult to address word meaning because it is a surface-level metric. |
| Approach: | They propose to use word embeddings, sentence-level tf-idf, and cosine similarity between two word embeds as features, weight, and the distance between two features as features. |
| Outcome: | The proposed metric can evaluate machine translation based on word meaning . it achieves highest correlation with human judgment among several representative metrics. |
Copied to clipboard
| Challenge: | Existing studies highlight inconsistencies between automated evaluation metrics and human expert assessments for patent claims. |
| Approach: | They propose a multi-dimensional evaluation method specifically designed for patent claims that incorporates features annotated by patent experts. |
| Outcome: | The proposed method achieves highest correlation with human expert evaluations across all assessment criteria across all tested metrics. |
Copied to clipboard
| Challenge: | Cognitive science and symbolic AI research suggest that event causality provides vital information for story understanding. |
| Approach: | They propose a method for event causality identification that leads to material improvements in story understanding. |
| Outcome: | The proposed method improves story understanding on the COPES dataset . it achieves 4.1-10.9% increase on Clip Accuracy and 4.2-13.5% increase on Sentence IoU . |
Copied to clipboard
| Challenge: | Existing studies on summary quality measure have shown that it correlates well with quality scores produced by human annotators. |
| Approach: | They propose to use a criterion that does not rely on human scores to judge summary quality . they propose to develop a method that can be used to determine the best measure from a family of measures . |
| Outcome: | The proposed measure could be used to determine the best summary quality measure from a family of measures. |
Copied to clipboard
| Challenge: | Existing graph-based methods only consider word relations or structure information, which neglect the correlation between them. |
| Approach: | They propose a Dual Graph network for Abstractive Sentence Summarization that captures word relations and structure information from sentences. |
| Outcome: | The proposed model outperforms state-of-the-art methods on two popular benchmark datasets. |
Copied to clipboard
| Challenge: | masked language models produce stronger correlations than auto-regressive models, but humans and models make different response selection mistakes. |
| Approach: | They propose to use spoken conversation as a model to measure human comprehension behaviour. |
| Outcome: | The proposed model outperforms the model which produces the strongest correlation with human responses. |
Copied to clipboard
| Challenge: | Existing studies have used the correlation information stored in samples for self-supervised learning, but they feed the training pairs in a random order without consideration of difficulty. |
| Approach: | They propose to inject curriculum learning into weakly supervised multimodal correlation learning by scoring and feeding pairs according to difficulty. |
| Outcome: | The proposed model achieves state-of-the-art on multimodal sentiment analysis without human annotation. |
Copied to clipboard
| Challenge: | Existing methods for estimating uncertainty using answer likelihoods or prompt-based confidence generation often suffer from overconfidence and confirmation biases. |
| Approach: | They propose to use Decompose and Compare Consistency (DeCC) to measure the reliability of a VLM's direct answer and indirect answers by decomposing the question into sub-questions and reasoning over the sub-answers. |
| Outcome: | Experiments on six vision-language tasks with three VLMs show that DeCC achieves better correlation with task accuracy compared to existing methods. |
Copied to clipboard
| Challenge: | Existing tools for dialogue evaluation do not generalize to unseen datasets and/or need a human-generated reference response during inference. |
| Approach: | They propose an unreferenced automated dialogue evaluation metric that uses large pre-trained language models to extract latent representations of utterances and leverages the temporal transitions that exist between them. |
| Outcome: | The proposed model achieves higher correlation with human annotations in an online setting, while not requiring true responses for comparison during inference. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are evolving rapidly and require manual evaluations. |
| Approach: | They propose an LLM-powered framework that automates the entire evaluation process using LLM agents. |
| Outcome: | The proposed framework shows a 92.14% correlation with human preferences, surpassing all previous expert-annotated benchmarks without any manual efforts. |
Copied to clipboard
| Challenge: | Existing literature has focused on pretrainer-based text-driven brain encoding models . however, few studies have explored the efficacy of task-specific learning of Transformers . |
| Approach: | They propose to use ten popular natural language processing tasks to learn Transformer representations for predicting brain responses. |
| Outcome: | The proposed model predicts brain activity across the whole brain. |
Copied to clipboard
| Challenge: | Existing models for large language models lack the ability to calibrate their outputs towards human preference. |
| Approach: | They propose a multi-stage, gradient-free approach to calibrate an LLM-based evaluator toward human preference. |
| Outcome: | The proposed approach improves correlation with expert evaluation on multiple text quality evaluation datasets. |
Copied to clipboard
| Challenge: | Instance-level difficulty analysis of evaluation data is a new field of research that focuses on leveraging instance difficulty in natural language processing. |
| Approach: | They conduct Instance-Level Difficulty Analysis of Evaluation data in a large-scale setup of 23 datasets and demonstrate its five novel applications. |
| Outcome: | The proposed model improves efficiency and accuracy, improves quality and improves Out-of-Domain performance. |
Copied to clipboard
| Challenge: | Existing open-source evaluation paradigms lack flexibility and performance . language model-based evaluation is cheap and scalable, but it is difficult to evaluate . |
| Approach: | They propose a language model-based evaluation paradigm that uses a scalar indicator of quality to assess LM outputs. |
| Outcome: | The proposed language model-based evaluation model is more powerful than its predecessor. |
Copied to clipboard
| Challenge: | Machine translation (MT) is currently evaluated in one of two ways: monolingually or trained crosslingually by building a supervised model to predict quality scores from human-labeled data. |
| Approach: | They propose an unsupervised model that directly compares the source and machine translated sentence using strong pretrained multilingual word and sentence representations. |
| Outcome: | The proposed model outperforms glass-box approaches to quality estimation that rely on a supervised model. |
Copied to clipboard
| Challenge: | Existing methods to narrate movies with no actors are difficult to implement in real situations . a new metric is proposed to provide the best correlation with human evaluation . |
| Approach: | They propose a large-scale Chinese movie benchmark to help visually impaired enjoy movies . they propose metric called Movie Narration Score (MNScore) which achieves best correlation with human evaluation. |
| Outcome: | The proposed method outperforms baselines and the existing methods. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) show impressive capabilities across visual–language tasks, but their capacity to evaluate artistic expression remains limited. |
| Approach: | They propose an attribute-specific multi-LoRA approach where each attribute corresponds to a distinct evaluation dimension in the scoring rubric. |
| Outcome: | The proposed approach increases correlation from 0.468 to 0.653 on Qwen2.5-VL-7B, with the largest gains on perceptual dimensions and narrowed gaps on higher-order attributes. |
Copied to clipboard
| Challenge: | Existing evaluation metrics are limited and can be easily portable to new languages. |
| Approach: | They propose a simple unsupervised metric and additional supervised metrics which rely on contextual word embeddings to encode the translation and reference sentences. |
| Outcome: | The proposed model outperforms existing metrics on the WMT 2017 dataset and is more accurate than existing models. |
Copied to clipboard
| Challenge: | ChatGPT and GPT-4 are popular as evaluation metric for complex generative tasks . however, they are not ready as human replacements due to significant limitations . |
| Approach: | They conduct extensive analysis to examine the stability and reliability of LLMs as automatic evaluators for abstractive summarization. |
| Outcome: | The proposed methods outperform the commonly used automatic metrics but are not ready for human evaluation due to significant limitations. |
Copied to clipboard
| Challenge: | Subword tokenization is a key part of most NLP pipelines, but little is known about why some combinations lead to improved downstream model performance. |
| Approach: | They propose that good tokenizers lead to efficient channel usage . they propose that an optimal encoding assigns extremely long codes to low-frequency subwords . |
| Outcome: | The proposed tokenizers have a very strong correlation with BLEU in machine translation . the proposed function can be used to improve model performance in the downstream task . |
Copied to clipboard
| Challenge: | Existing methods for machine translation evaluation do not require reference translations. |
| Approach: | They propose a method for machine translation evaluation which does not require reference translations. |
| Outcome: | The proposed method achieves highest correlation with human judgements on 9 out of 18 language pairs from the WMT19 benchmark for evaluation without references. |
Copied to clipboard
| Challenge: | Current evaluations of large language models rely on a single instruction template, overlooking models’ sensitivity to instruction style. |
| Approach: | They propose a multi-dimensional framework quantifying how instruction formulation affects model responses by transforming benchmark problems into multiple instruction styles. |
| Outcome: | The proposed framework reveals that instruction style can shift accuracy by 16.7% points. |
Copied to clipboard
| Challenge: | We compare attention functions in pre-trained language models to human eye fixation patterns during task-specific reading tasks. |
| Approach: | They compare attention functions in large-scale pre-trained language models to classical cognitive models of human attention by using a dataset with eye-tracking recordings of native speakers of English. |
| Outcome: | The proposed model is as predictive of human eye fixation patterns as classical cognitive models of human attention. |
Copied to clipboard
| Challenge: | Recent embedding-based evaluation metrics for text generation are based on measuring correlation with human evaluations on standard benchmarks. |
| Approach: | They examine the robustness of BERTScore, one of the most popular embedding-based metrics for text generation. |
| Outcome: | The embedding-based metrics that have the highest correlation with human evaluations on a standard benchmark can have the lowest correlation if the amount of input noise or unknown tokens increases. |
Copied to clipboard
| Challenge: | Existing automated RAG evaluation frameworks overlook important failure modes when using GPT-4 as a judge. |
| Approach: | They propose a novel pipeline to assess the calibration and discrimination capabilities of judge models by using a meta-evaluation benchmark of 144 unit tests to identify key failure modes. |
| Outcome: | The proposed pipeline improves on existing frameworks, while state-of-the-art open-source judges do not generalize to their proposed criteria. |
Copied to clipboard
| Challenge: | Reinforcement Learning (RL)-based document summarisation systems produce state-of-the-art performance in terms of ROUGE scores, but high summaries receive low human judgement. |
| Approach: | They propose to learn a reward function from human ratings on 2,500 summaries to generate human-appealing summary. |
| Outcome: | The proposed reward function can generate human-appealing summaries without reference summary input. |
Copied to clipboard
| Challenge: | Existing evaluation methods based on large language models (LLMs) are expensive and lack expertise due to limitations in human expertise. |
| Approach: | They propose an open-source automatic evaluation model with 13B parameters specifically engineered to measure the question-answering proficiency of medical LLMs. |
| Outcome: | The proposed model surpasses baselines in terms of correlation with human judgments. |
Copied to clipboard
| Challenge: | Existing studies on text simplification have focused on sentence simplification, but these metrics often underperform on longer texts. |
| Approach: | They propose to adapt existing sentence-level metrics for paragraph- or document-level simplification by incorporating a new approach to the evaluation of text simplification metrics. |
| Outcome: | The proposed approach outperforms existing sentence-level metrics in terms of correlation with human judgment and the sensitivity and robustness of various metrics to different types of errors produced by existing systems. |
Copied to clipboard
| Challenge: | Existing computational models capture prediction and reanalysis using Large Language Models (LLMs) and a statistical measure known as ‘surprisal’. |
| Approach: | They propose to extract structural information from Large Language Models and a statistical measure known as ‘surprisal’ to integrate it with their learnt statistics. |
| Outcome: | The proposed model achieved higher correlation with human reading times and better predicted the garden path effect and could distinguish between sentence types with different levels of difficulty. |
Copied to clipboard
| Challenge: | Neural machine translation models are often criticized for failures that happen without competency awareness. |
| Approach: | They propose a method that extends conventional NMT with a self-estimator to translate a source sentence and estimate its competency. |
| Outcome: | The proposed method performs on translation tasks intact and on quality estimation tasks better than existing methods. |
Copied to clipboard
| Challenge: | Existing methods to evaluate the quality of language generation do not provide explicit explanation of their verdicts. |
| Approach: | They propose a fine-grained explainable evaluation metric for text generation that harnesses human instruction and implicit knowledge of GPT-4 to fine-tune it. |
| Outcome: | The proposed model outperforms all other unsupervised metrics on translation, captioning, data-to-text, and commonsense generation tasks. |
Copied to clipboard
| Challenge: | Evaluating the quality of texts generated by language models has always been a challenging task in natural language processing (NLP). |
| Approach: | They propose a multidimensional comparative evaluation method based on instruction-following that combines relevance, factuality, and adherence with a concrete Chain-of-Thoughts process to enhance the accuracy of evaluations. |
| Outcome: | The proposed method outperforms existing methods in correlation with human evaluations on two NLG evaluation benchmarks. |
Copied to clipboard
| Challenge: | lexical overlap is a common evaluation metric for extractive summarization, but recent studies reveal its limitations. |
| Approach: | They propose a facet-aware evaluation setup for better assessment of information coverage in extractive summaries. |
| Outcome: | The proposed evaluation setup improves human correlation with extractive summarization datasets and improves comparative analysis. |
Copied to clipboard
| Challenge: | Existing automatic metrics do not capture errors in abstractive summarization models. |
| Approach: | They propose an automatic question answering metric for faithfulness that leverages recent advances in reading comprehension. |
| Outcome: | The proposed metric has significantly higher correlation with human faithfulness scores on highly abstracted summaries. |
Copied to clipboard
| Challenge: | Pretraining-based (PT) evaluation metrics are not effective for training grammatical error correction systems. |
| Approach: | They propose a pretraining-based GEC evaluation metric which only uses PT-based metrics to score the corrected parts of the system. |
| Outcome: | The proposed evaluation metric outperforms existing methods on a CoNLL14 evaluation task. |
Copied to clipboard
| Challenge: | X-Eval is a two-stage instruction tuning framework to evaluate text in both seen and unseen aspects customized by end users. |
| Approach: | They introduce a two-stage instruction tuning framework to evaluate text in both seen and unseen aspects customized by end users. |
| Outcome: | The proposed framework improves the model’s ability to follow evaluation instructions and enhances the learning stage to better assess text quality. |
Copied to clipboard
| Challenge: | Existing paradigm to fine-tune parameters of pre-trained language models poses problems in data-scarce and resource-limited scenarios. |
| Approach: | They propose a parameter-efficient fine-tuning method HiFi that fine-tails only the highly informative and strongly correlated attention heads for the specific task. |
| Outcome: | The proposed method obtains state-of-the-art over the prior benchmarks on the GLUE benchmark. |
Copied to clipboard
| Challenge: | N-gram-based evaluation metrics are unreliable due to low correlation to human judgments. |
| Approach: | They propose a metric that rewards correct details and penalizes incorrect ones. |
| Outcome: | The proposed metric matches the performance of open-source LLM-based metrics in correlation to human judgments while being far more efficient. |
Copied to clipboard
| Challenge: | Existing methods to assess fairness using prompts have low correlations between fairness metrics. |
| Approach: | They propose a method to enhance the correlation between fairness metrics by using pre-trained language models. |
| Outcome: | The proposed method improves the correlation between fairness metrics by using pre-trained language models. |
Copied to clipboard
| Challenge: | e-commerce has become a research hotspot for review helpfulness prediction . a new approach to help predict helpfulness of multimodal product reviews is proposed . |
| Approach: | They propose a machine learning task to identify helpfulness of multimodal product reviews . they use a probe-based strategy to enforce high attention weights on regions of greater significance . |
| Outcome: | The proposed model achieves state-of-the-art performance with lower memory consumption on two benchmark datasets with three categories. |
Copied to clipboard
| Challenge: | End-to-End speech-to speech translation is generally evaluated with text-based metrics . this means generated speech has to be automatically transcribed, making the evaluation dependent on ASR systems. |
| Approach: | They propose a text-free evaluation metric for end-to-end speech-tospeech translation, named BLASER, to avoid the dependency on automatic speech recognition systems. |
| Outcome: | The proposed metric avoids the dependency on automatic speech recognition systems by encoding generated speech segments into a shared embedding space. |
Copied to clipboard
| Challenge: | Existing metrics that compare the candidate with the human reference do not consider the context, resulting in poor correlation with human judgements. |
| Approach: | They propose a language model-aware metric that augments the human reference while considering the context to provide evaluation scores that correlate highly with human judgements. |
| Outcome: | The proposed metric achieves higher correlation with human reference judgements and differentiates well-formed candidates from adversarial samples to a larger degree. |
Copied to clipboard
| Challenge: | Existing methods for inferring abstractness of words and expressions without labeled data are limited and limited. |
| Approach: | They propose a weakly supervised approach for inferring the property of abstractness of words and expressions in the absence of labeled data. |
| Outcome: | The proposed approach obtains high correlation with human labels in the absence of labeled data. |
Copied to clipboard
| Challenge: | Automated summarization metrics are reliable but often poorly correlated with human judgment. |
| Approach: | They propose a semi-automatic to automatic summary evaluation metrics, following the Pyramid human evaluation method. |
| Outcome: | The proposed metrics are semi-automatic to automatic summary evaluation metrics, following the Pyramid human evaluation method. |
Copied to clipboard
| Challenge: | KG-to-Text models are prone to errors like Additions and Omissions, and few languages are taken into account since both train and test data are not readily available. |
| Approach: | They propose a multilingual evaluation framework that is reference-less . it allows estimating how much a KG-to-Text Model under- (omission) or over- (addition) generates. |
| Outcome: | The proposed evaluation framework outperforms prior reference-less metrics in correlation with human judgments and provides scores for precision and recall. |
Copied to clipboard
| Challenge: | Existing reference-less metrics are not optimized for manual evaluations of system outputs because no dataset exists for manual analysis. |
| Approach: | They propose a reference-less metric trained on manual evaluations of system outputs for grammatical error correction. |
| Outcome: | The proposed metric improves correlation with manual evaluation in system- and sentence-level meta-evaluation. |
Copied to clipboard
| Challenge: | Existing methods for evaluation of dialog systems are expensive and not scalable . a framework for estimating human evaluation scores is proposed to bridge this gap . |
| Approach: | They propose a framework for estimating human evaluation scores based on off-policy evaluation . they use language quality metrics for single-turn response generation given a fixed context . |
| Outcome: | The proposed framework outperforms existing methods in terms of correlation with human evaluation scores. |
Copied to clipboard
| Challenge: | Image captioning relies on reference-based automatic evaluations, but references are expensive to collect and comparing against multiple human-authored captions is insufficient. |
| Approach: | They propose a reference-free metric that can be used for automatic caption evaluation without references. |
| Outcome: | The proposed model outperforms existing metrics on image-text compatibility and a reference-augmented version achieves even higher correlation with human judgements. |
Copied to clipboard
| Challenge: | Recent-proposed evaluation metrics for large language models have a preference-bias . however, such metrics often lack interpretability and only offer a single score . |
| Approach: | They propose a metric that leverages the power of large language models to perform two sub-tasks: decomposing summaries into atomic content units and validating them against the source document. |
| Outcome: | The proposed metric improves faithfulness scores on three summarization evaluation benchmarks by 3% compared to the next-best metric. |
Copied to clipboard
| Challenge: | Automated evaluation of natural language generation tasks fails to focus on medical QA because of the diversity in medical terminology. |
| Approach: | They propose a new data structure, imap, to capture key information in questions and answers. |
| Outcome: | The proposed model outperforms state-of-the-art metrics in correlation with human scores. |
Copied to clipboard
| Challenge: | Existing evaluation methods for factual consistency in knowledge-grounded dialogues are unreliable and limit their applicability. |
| Approach: | They propose an automatic evaluation metric for factual consistency in knowledge-grounded dialogue using automatic question generation and question answering. |
| Outcome: | The proposed evaluation metric consistently shows higher correlation with human judgements. |
Copied to clipboard
| Challenge: | Existing benchmarks for reward models show a weak correlation with performance of optimized policies . existing benchmarks do not accurately assess the true capabilities of reward models . |
| Approach: | They explore how reward overoptimization captures how well a reward model aligns with human preferences and the dynamics of the learning signal it provides to the policy. |
| Outcome: | The proposed benchmarks show that reward overoptimization is a weak factor . the high correlation with degree of overoptimalization leads to lower correlation with downstream performance . |
Copied to clipboard
| Challenge: | Existing methods for evaluating factual consistency are primarily designed for short summaries of isolated code snippets. |
| Approach: | They propose a reference-free and fine-grained method for evaluating factual consistency in real-world code summaries. |
| Outcome: | The proposed method achieves highest correlation with human judgment among 13 baselines, improving 15-18% over the previous state-of-the-art. |
Copied to clipboard
| Challenge: | Existing methods for evaluating image description generation systems are subjective and expensive to scale. |
| Approach: | They propose a new image-aware metric for evaluating image description generation systems . it estimates the faithfulness of a generated caption with respect to the content of the actual image . |
| Outcome: | The proposed metric achieves high correlation with human judgments on two well-known datasets and is competitive with metrics that depend on and rely exclusively on human references. |
Copied to clipboard
| Challenge: | Graph-to-text (G2T) generation is an important task in natural language generation as it renders graphs accessible to non-technical users in downstream applications such as question answering. |
| Approach: | They propose a metric that correctly identifies factual faithfulness and uses it to determine if a triple is present in a generated text. |
| Outcome: | The proposed metric achieves highest correlation with human annotations on data correctness, data coverage, and relevance. |
Copied to clipboard
| Challenge: | Current evaluation practices of open domain dialogue systems are still highly dependent on human evaluation. |
| Approach: | They propose to use an annotated dataset to evaluate chatbots using large language models. |
| Outcome: | The proposed model improves over few-shot inferences on a GPT-3.5 generated dialogue dataset. |
Copied to clipboard
| Challenge: | a poor selection of an anchor can dramatically reduce correlation with human rankings . traditional reference-based metrics are often ill-suited for open-ended generation . |
| Approach: | They evaluate 22 different anchors on a Arena-Hard-v2.0 dataset and quantify the effect size of anchor selection. |
| Outcome: | The proposed model is better or worse than all other models, but it is rarely indicative of the relative ranking of the models. |
Copied to clipboard
| Challenge: | Existing benchmarks measure the correlation with human judgements of faithfulness on model-generated summaries, but they are insufficient for diagnosing whether metrics are consistent, effective on human-written texts, and sensitive to different error types. |
| Approach: | They propose to use unfaithful minimal pairs to measure the consistency of automatic faithfulness metrics by comparing human-written summary pairs with a dataset of 889 human-writing, minimally different summary pairs. |
| Outcome: | The proposed benchmarks show that the most discriminative metrics tend not to be the most consistent, and that the best performing metrics are sensitive to errors. |
Copied to clipboard
| Challenge: | Image captioning has been a challenge for vision-language researchers for decades . current VLMs focus on tasks like visual question answering (YA) but image captioning is not as advanced as expected. |
| Approach: | They evaluate VLMs' performance on image captioning using human annotations . they find that some metrics show high caption-level agreement with humans . |
| Outcome: | The proposed model outperforms open-source models on image captioning . it achieves 93.4% correlation with human rankings at $4 per test . |
Copied to clipboard
| Challenge: | Existing evaluation metrics for automated audio captioning only provide an overall score . current evaluation checklists are inadequate to characterize the nuanced differences . |
| Approach: | They propose an explainable and multi-factor audio captioning evaluation paradigm . they define sound event, source, attribute and relation as four factors tailored for the audio description . |
| Outcome: | The proposed evaluation paradigm improves the quality of audio captions . it can detect mismatches and align with human perception, the authors show . |
Copied to clipboard
| Challenge: | Existing evaluation metrics conflate simplicity with correlated attributes such as fluency or meaning preservation. |
| Approach: | They propose a new learning evaluation metric that focuses on simplicity outperforming most existing metrics in terms of correlation with human judgements. |
| Outcome: | The proposed metric outperforms most existing metrics in terms of correlation with human judgements. |
Copied to clipboard
| Challenge: | linguistic models have a higher correlation with human ground truth ratings than labeled data . word vectors have often been evaluated on standard word relatedness benchmarks . |
| Approach: | They propose to use unsupervised, supervised, and finally supervised methods to extract emotional associations from pretrained vectors and models. |
| Outcome: | The proposed method shows higher correlation with ground truth ratings than state-of-the-art lexicons based on labeled data. |
Copied to clipboard
| Challenge: | Emphasis is a crucial component in human communication, which indicates speaker’s intention and implication beyond pure text in dialogue. |
| Approach: | They propose a benchmark dataset with annotated dialogue samples capturing the implications of emphasis. |
| Outcome: | The proposed evaluation pipeline achieves high correlation with human scoring and commercial LLMs perform better than open-source LLM. |
Copied to clipboard
| Challenge: | Current evaluation of neural machine translation systems is limited by one best hypothesis and search errors brought by heuristic decoding algorithms. |
| Approach: | They propose a new evaluation protocol which defines model errors with model’s ranking capability over hypothesis space and Monte Carlo sampling evaluation to tackle the problem of exponentially large space. |
| Outcome: | The proposed evaluation protocol is consistent with what is currently used in the field and is consistent to what is being proposed. |
Copied to clipboard
| Challenge: | Summarization evaluation approaches have relied on ROUGE for summarization, but they fall short of human evaluations. |
| Approach: | They propose a new approach to evaluate summaries by leveraging retrieval techniques . they use a dual-encoder retrieval setup to train a retrieval task . |
| Outcome: | The proposed method outperforms existing methods on two document summarization benchmarks and a long document summmarization test. |
Copied to clipboard
| Challenge: | Existing approaches to evaluate open domain dialogues have a one-to-many problem . existing approaches lack commonsense reasoning biases and perform poorly in domain-specific scenarios. |
| Approach: | They propose a framework that leverages both a small, specialised model and LLMs for the evaluation of open-domain dialogues. |
| Outcome: | The proposed framework achieves state-of-the-art performance in both classification and evaluation tasks and exhibits better correlation with human judgements. |
Copied to clipboard
| Challenge: | n-gram-based metrics are considered indicative (even if imperfect) of human evaluation for English, but their suitability for other languages remains unclear. |
| Approach: | They systematically assess evaluation metrics for generation for languages and tasks using n-gram-based and neural-based metrics. |
| Outcome: | The proposed evaluation suite is based on eight languages from four typological families and shows that it is sensitivity to the language type at hand. |
Copied to clipboard
| Challenge: | LLM judges have gained popularity as an inexpensive and performant substitute for human evaluation. |
| Approach: | They revisit meta-evaluations of LLM evaluators under a setting that more closely aligns with practice by examining evaluers’ ability to distinguish test system pairs that are closer in capability. |
| Outcome: | The proposed meta-evaluation setting is significantly different from the use of human evaluations. |
Copied to clipboard
| Challenge: | a large amount of insight into human language processing can be gleaned by studying word-by-word processing difficulty. |
| Approach: | They extend the study by examining eyetracking corpora of seven languages . they find evidence for superlinearity in some languages, but highly sensitive to language models . |
| Outcome: | The study extends existing studies on english to Danish, Dutch, English, German, Japanese, Mandarin, and Russian. |
Copied to clipboard
| Challenge: | Recent studies have shown that MT metrics return assessments as scalar scores that are difficult to interpret, posing a challenge to making informed design choices. |
| Approach: | They propose an interpretable evaluation framework that evaluates MT metrics in two scenarios that serve as proxies for filtering and translation re-ranking use cases. |
| Outcome: | The proposed framework offers clearer insights than correlation with human judgments. |
Copied to clipboard
| Challenge: | Existing word embedding methods overlook phonetic information that is crucial for many tasks. |
| Approach: | They propose three methods that use articulatory features to build phonetically informed word embeddings. |
| Outcome: | The proposed methods improve word retrieval and correlation with sound similarity and on rhyme and cognate detection tasks. |
Copied to clipboard
| Challenge: | The paper presents the design and construction of a time-aligned multimodal dataset for reading research, including multiple time-aligned temporal signals elicited with four experimental trials of connected text reading by both child and adult readers. |
| Approach: | They propose to use a time-stamped multimodal dataset to analyze time-aligned temporal signals elicited by connected text reading by both child and adult readers. |
| Outcome: | The proposed dataset includes multiple time-aligned temporal signals elicited with four experimental trials of connected text reading by both child and adult readers. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated near-human performance in summarization tasks based on traditional metrics such as ROUGE and BERTScore . however, these metrics do not adequately capture critical aspects of summarizing quality, such as factual accuracy, especially for long narratives. |
| Approach: | They propose a framework that evaluates and refines factuality in narrative summarization by leveraging a Character Knowledge Graph extracted from input narrative. |
| Outcome: | The proposed framework evaluates factuality and provides actionable guidance for refinement. |
Copied to clipboard
| Challenge: | Recent studies show that LLM-based agents struggle to perform in zero-shot scenarios. |
| Approach: | They propose a framework to quantify the behavior gap between AI agents and human experts . they propose to examine discrepancies in dialog acts, tool usage, and knowledge utilization . |
| Outcome: | The proposed framework measures the behavior gap between AI agents and human experts on task-oriented dialogs. |
Copied to clipboard
| Challenge: | State-of-the-art trainable machine translation evaluation metrics rely on large encoders . this makes them computationally expensive and inaccessible to researchers with limited resources. |
| Approach: | They propose a method to extract knowledge stored in large encoders and a pipeline for efficient black-box distillation. |
| Outcome: | The proposed model surpasses COMET-22 and BLEURT-20 on the WMT22 dataset by 6.4%. |
Copied to clipboard
| Challenge: | Existing reference-free automatic grammatical error correction methods do not correlate with human evaluation. |
| Approach: | They propose a reference-free automatic grammatical error correction evaluation method with enhanced gramma-ed capabilities. |
| Outcome: | The proposed method achieves highest correlation with human evaluations on a meta-evaluation dataset. |
Copied to clipboard
| Challenge: | Existing LLMs focus on responding to specific arguments while neglecting objective assessments such as authenticity and logical validity. |
| Approach: | They propose a multi-dimensional evaluation system and an optimized debating framework . they propose to use coT reasoning enhancement, web-based Retrieval Augmented Generation to optimize across various dimensions. |
| Outcome: | The proposed framework outperforms baseline models in argument quality assessment and debate process simulation by 57%. |
Copied to clipboard
| Challenge: | Existing video-to-text summarization evaluation methods depend heavily on human-written reference summaries. |
| Approach: | They propose a reference-free metric evaluating candidate summaries directly against source videos through multimodal question answering. |
| Outcome: | The proposed metric assesses candidate summaries directly against source videos through multimodal question answering. |
Copied to clipboard
| Challenge: | Reference-free evaluation metrics for grammatical error correction have high correlation with human judgments, but they are not designed to evaluate adversarial systems that aim to obtain unjustifiably high scores. |
| Approach: | They propose adversarial attack strategies for four reference-free metrics . they propose SOME, Scribendi, IMPARA, and LLM-based metrics based on these metrics a . |
| Outcome: | The proposed attacks outperform the current state-of-the-art for four reference-free metrics . |
Copied to clipboard
| Challenge: | Existing evaluation metrics for literature prioritize mechanical accuracy over artistic expression . this bias could result in an irreversible decline in translation quality and cultural authenticity . |
| Approach: | They propose a novel, reference-free, LLM-based question-answering framework for literary translation evaluation. |
| Outcome: | a novel, reference-free, LLM-based question-answering framework is developed for literary translation evaluation. |
Copied to clipboard
| Challenge: | Existing metrics for caption evaluation lack factual accuracy and limited context handling . VC-Inspector provides reproducible, fact-aware alternative that aligns closely with human judgments. |
| Approach: | They propose a lightweight, open-source large multimodal model for reference-free evaluation of video captions with a focus on factual accuracy. |
| Outcome: | Experiments show that VC-Inspector can generalize across diverse domains and improve on existing metrics. |
Copied to clipboard
| Challenge: | Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban aerial spaces remain to be explored. |
| Approach: | They propose a benchmark to evaluate whether large multimodal models can process continuous first-person visual observations like humans. |
| Outcome: | The proposed model can process first-person visual observations like humans, enabling recall, perception, reasoning, and navigation. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown promise in generating visualizations from natural language, but lack of comprehensive benchmarks limits their capabilities. |
| Approach: | They propose a framework that jointly refines the textual answer and visualization code to improve GPT-4o's pass rate from 26% to 42% over direct approach. |
| Outcome: | The proposed framework increases GPT-4o’s pass rate from 26% to 42% over the direct approach and improves chart quality. |
Copied to clipboard
| Challenge: | closed-ended question-based benchmarks struggle with saturation as newer models emerge . crowd-sourced leaderboards rely on costly and slow human judges . |
| Approach: | They propose a framework that leverages collective intelligence from all large language models to evaluate each other. |
| Outcome: | a new framework enables a democratic, pairwise evaluation of all large language models . it achieves 97% correlation with human judgements, while significantly reducing the cost. |
Copied to clipboard
| Challenge: | a dataset of Tunisian Arabic instructions and prompts is used to evaluate LLMs' ability to understand and generate responses in Tunisia . we assess the quality, correctness, relevance, and dialectal adherence of LLM responses . |
| Approach: | They propose a benchmark for evaluating the capabilities of large language models in Tunisian Arabic . they use a dataset of Tunisia Arabic instructions and prompts to evaluate their models . |
| Outcome: | The proposed model can judge quality, correctness, relevance, and dialectal adherence . the model can also generate a leaderboard for the Tunisian Arabic language . |
Copied to clipboard
| Challenge: | Existing LAM benchmarks with thousands of examples create substantial computational barriers. |
| Approach: | They examine whether subsets can reliably evaluate large audio models . they find that subset of 50 examples can achieve over 0.93 Pearson correlation with full benchmark . |
| Outcome: | The proposed method outperforms the full benchmark and subset selection methods. |
Copied to clipboard
| Challenge: | Existing evaluation metrics fail to evaluate factual correctness in procedural video captions . Existing metrics rely on lexical overlap or holistic semantic similarity, but miss role-specific omissions resulting in hallucinations . |
| Approach: | They propose a role-aware, fact-level evaluation framework that distinguishes conceptual facts from contextual facts. |
| Outcome: | Experiments show that state-of-the-art captioning models produce fluent but incomplete descriptions with systematic errors. |
Copied to clipboard
| Challenge: | Personalized MGT detection remains largely underexplored due to personalization challenges . large language models (LLMs) can imitate personal writing styles, but they can generate fake news and misinformation. |
| Approach: | They propose a benchmark to evaluate detector robustness under personalization . they attribute this limitation to a feature-inversion trap that flips the effect in personalized contexts . |
| Outcome: | The proposed framework predicts detector robustness under personalization with an 85% correlation to actual results. |
Copied to clipboard
| Challenge: | Existing metrics for evaluating the quality of tables generated by large language models flatten tables into text, ignoring structure or relying on fixed references that limit generalization. |
| Approach: | They propose a reference-less framework for evaluating tabular generation via graph-based reasoning . tabReX converts source text and generated tables into canonical knowledge graphs . |
| Outcome: | The proposed framework provides a high correlation with expert rankings and stable under harder perturbations. |