Papers with correlation

124 papers
Learning-based Composite Metrics for Improved Caption Evaluation (P18-3)

Copied to clipboard

Challenge: Existing image captioning metrics focus on linguistic aspects and do not match human judgements at sentence-level.
Approach: They propose to incorporate lexical and semantic metrics as features to capture adequacy and fluency of captions at different linguistic levels.
Outcome: The proposed framework captures adequacy and fluency of captions at different linguistic levels.
JUDGEBERT: Assessing Legal Meaning Preservation Between Sentences (2025.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics for text simplification focus on only one dimension: fluency, simplicity and meaning preservation.
Approach: They introduce a dataset to assess legal meaning preservation between two legal texts . they also introduce sanity checks for two identical sentences .
Outcome: The proposed metric shows superior correlation with human judgment compared to existing metrics.
BLEU might be Guilty but References are not Innocent (2020.emnlp-main)

Copied to clipboard

Challenge: Using a method to collect references and compare their value with human evaluations, we show that multi-reference BLEU does not improve the correlation for high quality output.
Approach: They propose a method to compare the quality of automated metrics by analyzing references and comparing them with human evaluations.
Outcome: The proposed method improves correlation with all modern evaluation metrics including embedding-based methods.
Robust Hate Speech Detection via Mitigating Spurious Correlations (2022.aacl-short)

Copied to clipboard

Challenge: a novel hate speech detection model can be used to detect word- and character-level adversarial attacks . existing adversarials assume that attackers replace the target words with other names to evade detection .
Approach: They propose a robust hate speech detection model that can defend against adversarial attacks . they describe the process of hate speech recognition by a causal graph and a regularized entropy loss function to quantify spurious correlation .
Outcome: The proposed model can defend against word- and character-level adversarial attacks.
Visio-Linguistic Brain Encoding (2022.coling-1)

Copied to clipboard

Challenge: Existing studies have failed to explore co-attentive multi-modal modeling for visual and text reasoning.
Approach: They propose to use image and multi-modal Transformers to reconstruct fMRI brain activity . they use two popular datasets to study visual and text reasoning .
Outcome: The proposed model outperforms existing models on two popular datasets . the results raise the question whether visual processing is affected implicitly by linguistic processing .
Barriers to Effective Evaluation of Simultaneous Interpretation (2024.findings-eacl)

Copied to clipboard

Challenge: Existing studies have relied on out-of-the-box machine translation metrics to evaluate interpretation data, but they do not account for human judgments of interpretation quality.
Approach: They propose to use machine translation metrics to evaluate human interpretations to address potential barriers to disfluency, summarization, paraphrasing and segmentation.
Outcome: The proposed model achieves better correlation with human judgments than state-of-the-art metrics.
CLIReval: Evaluating Machine Translation as a Cross-Lingual Information Retrieval Task (2020.acl-demos)

Copied to clipboard

Challenge: evaluating machine translation (MT) with cross-lingual information retrieval is relatively time-consuming and subjective.
Approach: They propose a toolkit that evaluates machine translation with a proxy task of cross-lingual information retrieval.
Outcome: The proposed toolkit is based on the "metrics shared task" of WMT2019.
Efficient Online Scalar Annotation with Bounded Support (P18-1)

Copied to clipboard

Challenge: Existing methods for efficiently eliciting scalar annotations for dataset construction and system quality estimation by human judgments are not shown.
Approach: They propose a method for efficiently eliciting scalar annotations by human judgments.
Outcome: The proposed method leads to increased correlation with ground truth, suggesting it is an improved mechanism for dataset creation and manual system evaluation.
GREEN: Generative Radiology Report Evaluation and Error Notation (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing automated evaluation metrics fail to consider factual correctness or are limited in their interpretability.
Approach: They propose a radiology report evaluation metric that leverages natural language understanding of language models to identify and explain clinically significant errors.
Outcome: The proposed method demonstrates higher correlation with expert error counts and higher alignment with expert preferences when compared to previous methods.
Multi-lingual Common Semantic Space Construction via Cluster-consistent Word Embedding (D18-1)

Copied to clipboard

Challenge: a new approach to multilingual word embedding is needed to achieve this goal . a multilingual common semantic space is a language-agnostic semantic continuous space .
Approach: They propose a multilingual common semantic space where words from multiple languages are mapped into a shared space so that resources and knowledge can be shared across languages.
Outcome: The proposed approach achieves 14.6% absolute F-score gain over state-of-the-art methods on cross-lingual direct transfer.
Unsupervised Learning of KB Queries in Task-Oriented Dialogs (2021.tacl-1)

Copied to clipboard

Challenge: Existing approaches require dialog datasets to explicitly annotate knowledge base (KB) queries.
Approach: They propose a pipelined approach to predict when to make a KB query and train the dialog agent without explicit annotation.
Outcome: The proposed approach predicts when to make a KB query, then predicts a query at the predicted position and uses the results in subsequent dialog.
GUM-SAGE: A Novel Dataset and Approach for Graded Entity Salience Prediction (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for graded entity salience are subjective but lack consistency.
Approach: They propose a method for graded entity salience that combines subjective judgments and summarization-based methods that define saliency as mention-worthiness in a summary.
Outcome: The proposed approach outperforms existing methods and shows stronger correlation with human summaries and alignments.
Ev2R: Evaluating Evidence Retrieval in Automated Fact-Checking (2026.tacl-1)

Copied to clipboard

Challenge: Current methods for automated fact-checking rely on relying on other evaluation metrics and closed knowledge sources.
Approach: They propose a method which combines evidence evaluation with verdict-level proxy scoring.
Outcome: The proposed method outperforms existing methods in accuracy and robustness against human ratings and adversarial tests.
An Error Analysis Framework for Shallow Surface Realization (2021.tacl-1)

Copied to clipboard

Challenge: BLEU and METEOR metrics fail to provide information on which linguistic factors impact performance of natural language generation models.
Approach: They propose a framework for error analysis which permits identifying which features of the input affect the models’ results.
Outcome: The proposed framework improves the performance of 174 system runs submitted to the Multilingual SR shared tasks.
Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles (2024.naacl-long)

Copied to clipboard

Challenge: Existing studies on multi-document summarization focus on collating information that all sources agree upon, but the task of summarizing diverse information remains underexplored.
Approach: They propose a task of summarizing diverse information encountered in multiple news articles encompassing the same event using a dataset curated by a large language model.
Outcome: The proposed task aims to summarize diverse information in multiple news articles encompassing the same event . the proposed task is difficult due to its limited coverage and verbosity biases .
Learning to Compare Financial Reports for Financial Forecasting (2024.findings-eacl)

Copied to clipboard

Challenge: Public companies in the US are required to publish annual reports that contain over 25,000 words across all sections and a high percentage of boilerplate content that does not change much year-to-year.
Approach: They propose to model complex, cross-document relationships between financial reports using paired financial reports.
Outcome: The proposed model can predict company risk and correlation from financial reports . the proposed model is able to recognize complex, nuanced relationships with complex signals .
Unsupervised Quality Estimation for Neural Machine Translation (2020.tacl-1)

Copied to clipboard

Challenge: Existing approaches require large amounts of expert annotated data, computation, and time for training.
Approach: They propose an unsupervised approach to QE where no training is required . they use a dataset that enables work on both black-box and glass-box approaches .
Outcome: The proposed approach rivals state-of-the-art supervised QE models in terms of correlation with human judgments of quality.
Open-Domain Dialog Evaluation Using Follow-Ups Likelihood (2022.coling-1)

Copied to clipboard

Challenge: Existing methods do not correlate strongly with human annotations.
Approach: They propose a method that measures the probability that a language model will continue the conversation with a fixed set of follow-ups.
Outcome: The proposed method achieves the highest correlation with human evaluations when compared against twelve existing methods.
Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics? (2025.naacl-srw)

Copied to clipboard

Challenge: Text style transfer (TST) is a multidimensional task requiring the assessment of style transfer accuracy, content preservation, and naturalness.
Approach: They propose to use text style transfer metrics to evaluate outputs of text editors . they also investigate the potential of large language models as tools for TST evaluation .
Outcome: The proposed methods provide better insights than existing metrics, the authors show . their meta-evaluation through correlation with hu-man judgments shows they are effective .
Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies on multilingual image captioning have been hampered by a lack of high-quality evaluation datasets.
Approach: They present a dataset of 3600 images annotated with human-generated captions in 36 languages.
Outcome: The proposed dataset shows that it is feasible to build multilingual image captioning models trained on machine-translated data.
Test-time Adaptation for Machine Translation Evaluation by Uncertainty Minimization (2023.acl-long)

Copied to clipboard

Challenge: evaluators of machine translation systems often use text-based metrics to evaluate performance . however, these metrics lack semantic-level information and exhibit poor correlation with human ratings . authors propose a method to reduce inference bias of neural metrics in out-of-distribution data .
Approach: They propose to reduce inference bias by using uncertainty estimation, test-time adaptation, and inference to reduce model uncertainty.
Outcome: The proposed method reduces model uncertainty and improves correlation performance across models.
Improving Dialog Evaluation with a Multi-reference Adversarial Dataset and Large Scale Pretraining (2020.tacl-1)

Copied to clipboard

Challenge: Existing models for dialog evaluation are trained using a single relevant response and multiple random negatives.
Approach: They propose a dataset to test whether model-based dialog evaluation metrics can be used to train models . they propose n-gram based metrics and embedding based ones to be used for model-driven evaluation .
Outcome: The proposed model outperforms existing models on a reddit dataset on relevant responses and adversarial responses.
Estimating User Communication Styles for Spoken Dialogue Systems (2020.lrec-1)

Copied to clipboard

Challenge: a neural network estimation system for spoken dialogues can be used to estimate the communication style of a user's interaction, but this is rarely implemented in a live system.
Approach: They propose a neural network approach to estimate the communication style of spoken interaction, namely elaborateness and directness.
Outcome: The proposed method can estimate the elaborateness and directness of spoken interaction and improve the results with additional linguistic features.
Identifying Weaknesses in Machine Translation Metrics Through Minimum Bayes Risk Decoding: A Case Study for COMET (2022.aacl-main)

Copied to clipboard

Challenge: Neural metrics have a high correlation with human judgements but they are hard to eliminate due to their "black box" nature.
Approach: They propose to use minimum bayes risk decoding to explore and quantify weaknesses in COMET models.
Outcome: The proposed model is not sensitive enough to discrepancies in numbers and named entities, and is hard to remove by training on additional synthetic data.
What just happened? Evaluating retrofitted distributional word vectors (N19-1)

Copied to clipboard

Challenge: Recent work has attempted to enhance vector space representations using information from structured semantic resources.
Approach: They propose a root-mean-square error evaluation metric to evaluate the utility of different lexical resources for retrofitting.
Outcome: The proposed method improves word similarity performance by using root-mean-square error (RMSE) and root-macro-error (RMME) metric.
Multi-Hypothesis Machine Translation Evaluation (2020.acl-main)

Copied to clipboard

Challenge: Reliably evaluating Machine Translation (MT) through automated metrics is a long-standing problem.
Approach: They propose to use MT models to generate multiple diverse translations and use them as surrogates to reference translations to obtain a quantification of translation variability.
Outcome: The proposed approach improves correlation with human judgements of quality by 15%.
Patent-CR: A Dataset for Patent Claim Revision (2025.naacl-long)

Copied to clipboard

Challenge: Patent-CR is the first dataset created for the patent claim revision task in English.
Approach: They propose to create a dataset for the patent claim revision task in English that includes both initial patent applications rejected by examiners and the final granted versions.
Outcome: The proposed dataset includes both initial patent applications rejected by examiners and the final granted versions.
The Feasibility of Embedding Based Automatic Evaluation for Single Document Summarization (D19-1)

Copied to clipboard

Challenge: Existing evaluation methods for summarization systems measure semantic overlap between a system summary and a human reference on word-string level.
Approach: They propose to use distributed representations to evaluate system summary and human reference on word-string level.
Outcome: The proposed representations outperform ROUGE on recent corpora but are less good on test data used in previous studies.
Revisiting Automatic Evaluation of Extractive Summarization Task: Can We Do Better than ROUGE? (2022.findings-acl)

Copied to clipboard

Challenge: Existing methods to evaluate text summarization tasks using ROUGE have been criticized for lack of semantic understanding.
Approach: They propose a semantic-aware metric for extractive summarization task that is semantic-based . they use CNN/DailyMail dataset to study the new metric .
Outcome: The proposed metric is semantic-aware and shows higher correlation with human judgement and yields a large number of disagreements with the original ROUGE metric.
Multi-Task Reinforcement Learning for Enhanced Multimodal LLM-as-a-Judge (2026.acl-industry)

Copied to clipboard

Challenge: Existing MLLMs are optimized for single-task scenarios and struggle to generalize to diverse contexts.
Approach: They propose a framework that integrates multitask reinforcement learning and generalization capabilities of MLLMs to optimize the judge model across multiple tasks.
Outcome: The proposed framework outperforms baseline models in judgment consistency and correlation with human preferences.
Towards a Unified Multi-Dimensional Evaluator for Text Generation (2022.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation frameworks for natural language generation are dominated by similarity-based metrics.
Approach: They propose a multi-dimensional evaluator for natural language generation that integrates multiple dimensions into one evaluer.
Outcome: The proposed evaluator improves on three typical NLG tasks and improves with external knowledge.
Modeling Event Plausibility with Consistent Conceptual Abstraction (2021.naacl-main)

Copied to clipboard

Challenge: Understanding natural language requires common sense, one aspect of which is the ability to discern the plausibility of events.
Approach: They propose a method of forcing model consistency that improves correlation with human plausibility judgements.
Outcome: The proposed method improves correlation with human plausibility judgements.
KPQA: A Metric for Generative Question Answering Using Keyphrase Weights (2021.naacl-main)

Copied to clipboard

Challenge: Existing n-gram similarity metrics fail to discriminate the incorrect answers due to the free-form of the answer.
Approach: They propose a new metric that assigns different weights to each token via keyphrase prediction to judge the correctness of GenQA.
Outcome: The proposed metric has a significantly higher correlation with human judgments than existing metrics in various datasets.
SMURF: SeMantic and linguistic UndeRstanding Fusion for Caption Evaluation via Typicality Analysis (2021.acl-long)

Copied to clipboard

Challenge: Visual captioning is an open-ended area for evaluation, requiring specialized training to improve human-correlation.
Approach: They propose a new evaluation framework rooted in information theory . they propose metric SPURTS and metric SMURF to measure fluency .
Outcome: The proposed metrics achieve state-of-the-art correlation with human judgment compared with other evaluation metrics.
InfoMetIC: An Informative Metric for Reference-free Image Caption Evaluation (2023.acl-long)

Copied to clipboard

Challenge: Existing image captioning metrics provide a single score to measure caption qualities, which are less explainable and informative.
Approach: They propose an Informative Metric for Reference-free Image Caption evaluation to support this feedback . they propose to provide a text precision score, a vision recall score and an overall quality score .
Outcome: The proposed method improves on existing metrics on multiple benchmarks and compares coarse-grained scores with human judgements.
Word Embedding-Based Automatic MT Evaluation Metric using Word Position Information (N19-1)

Copied to clipboard

Challenge: Existing evaluation metrics for machine translation are difficult to address word meaning because it is a surface-level metric.
Approach: They propose to use word embeddings, sentence-level tf-idf, and cosine similarity between two word embeds as features, weight, and the distance between two features as features.
Outcome: The proposed metric can evaluate machine translation based on word meaning . it achieves highest correlation with human judgment among several representative metrics.
Towards Better Evaluation for Generated Patent Claims (2025.acl-long)

Copied to clipboard

Challenge: Existing studies highlight inconsistencies between automated evaluation metrics and human expert assessments for patent claims.
Approach: They propose a multi-dimensional evaluation method specifically designed for patent claims that incorporates features annotated by patent experts.
Outcome: The proposed method achieves highest correlation with human expert evaluations across all assessment criteria across all tested metrics.
Event Causality Is Key to Computational Story Understanding (2024.naacl-long)

Copied to clipboard

Challenge: Cognitive science and symbolic AI research suggest that event causality provides vital information for story understanding.
Approach: They propose a method for event causality identification that leads to material improvements in story understanding.
Outcome: The proposed method improves story understanding on the COPES dataset . it achieves 4.1-10.9% increase on Clip Accuracy and 4.2-13.5% increase on Sentence IoU .
Is Human Scoring the Best Criteria for Summary Evaluation? (2021.findings-acl)

Copied to clipboard

Challenge: Existing studies on summary quality measure have shown that it correlates well with quality scores produced by human annotators.
Approach: They propose to use a criterion that does not rely on human scores to judge summary quality . they propose to develop a method that can be used to determine the best measure from a family of measures .
Outcome: The proposed measure could be used to determine the best summary quality measure from a family of measures.
Integrating Semantic Scenario and Word Relations for Abstractive Sentence Summarization (2021.emnlp-main)

Copied to clipboard

Challenge: Existing graph-based methods only consider word relations or structure information, which neglect the correlation between them.
Approach: They propose a Dual Graph network for Abstractive Sentence Summarization that captures word relations and structure information from sentences.
Outcome: The proposed model outperforms state-of-the-art methods on two popular benchmark datasets.
Do dialogue representations align with perception? An empirical study (2023.eacl-main)

Copied to clipboard

Challenge: masked language models produce stronger correlations than auto-regressive models, but humans and models make different response selection mistakes.
Approach: They propose to use spoken conversation as a model to measure human comprehension behaviour.
Outcome: The proposed model outperforms the model which produces the strongest correlation with human responses.
Curriculum Learning Meets Weakly Supervised Multimodal Correlation Learning (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies have used the correlation information stored in samples for self-supervised learning, but they feed the training pairs in a random order without consideration of difficulty.
Approach: They propose to inject curriculum learning into weakly supervised multimodal correlation learning by scoring and feeding pairs according to difficulty.
Outcome: The proposed model achieves state-of-the-art on multimodal sentiment analysis without human annotation.
Decompose and Compare Consistency: Measuring VLMs’ Answer Reliability via Task-Decomposition Consistency Comparison (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for estimating uncertainty using answer likelihoods or prompt-based confidence generation often suffer from overconfidence and confirmation biases.
Approach: They propose to use Decompose and Compare Consistency (DeCC) to measure the reliability of a VLM's direct answer and indirect answers by decomposing the question into sub-questions and reasoning over the sub-answers.
Outcome: Experiments on six vision-language tasks with three VLMs show that DeCC achieves better correlation with task accuracy compared to existing methods.
Learning an Unreferenced Metric for Online Dialogue Evaluation (2020.acl-main)

Copied to clipboard

Challenge: Existing tools for dialogue evaluation do not generalize to unseen datasets and/or need a human-generated reference response during inference.
Approach: They propose an unreferenced automated dialogue evaluation metric that uses large pre-trained language models to extract latent representations of utterances and leverages the temporal transitions that exist between them.
Outcome: The proposed model achieves higher correlation with human annotations in an online setting, while not requiring true responses for comparison during inference.
Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are evolving rapidly and require manual evaluations.
Approach: They propose an LLM-powered framework that automates the entire evaluation process using LLM agents.
Outcome: The proposed framework shows a 92.14% correlation with human preferences, surpassing all previous expert-annotated benchmarks without any manual efforts.
Neural Language Taskonomy: Which NLP Tasks are the most Predictive of fMRI Brain Activity? (2022.naacl-main)

Copied to clipboard

Challenge: Existing literature has focused on pretrainer-based text-driven brain encoding models . however, few studies have explored the efficacy of task-specific learning of Transformers .
Approach: They propose to use ten popular natural language processing tasks to learn Transformer representations for predicting brain responses.
Outcome: The proposed model predicts brain activity across the whole brain.
Calibrating LLM-Based Evaluator (2024.lrec-main)

Copied to clipboard

Challenge: Existing models for large language models lack the ability to calibrate their outputs towards human preference.
Approach: They propose a multi-stage, gradient-free approach to calibrate an LLM-based evaluator toward human preference.
Outcome: The proposed approach improves correlation with expert evaluation on multiple text quality evaluation datasets.
ILDAE: Instance-Level Difficulty Analysis of Evaluation Data (2022.acl-long)

Copied to clipboard

Challenge: Instance-level difficulty analysis of evaluation data is a new field of research that focuses on leveraging instance difficulty in natural language processing.
Approach: They conduct Instance-Level Difficulty Analysis of Evaluation data in a large-scale setup of 23 datasets and demonstrate its five novel applications.
Outcome: The proposed model improves efficiency and accuracy, improves quality and improves Out-of-Domain performance.
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing open-source evaluation paradigms lack flexibility and performance . language model-based evaluation is cheap and scalable, but it is difficult to evaluate .
Approach: They propose a language model-based evaluation paradigm that uses a scalar indicator of quality to assess LM outputs.
Outcome: The proposed language model-based evaluation model is more powerful than its predecessor.
SentSim: Crosslingual Semantic Evaluation of Machine Translation (2021.naacl-main)

Copied to clipboard

Challenge: Machine translation (MT) is currently evaluated in one of two ways: monolingually or trained crosslingually by building a supervised model to predict quality scores from human-labeled data.
Approach: They propose an unsupervised model that directly compares the source and machine translated sentence using strong pretrained multilingual word and sentence representations.
Outcome: The proposed model outperforms glass-box approaches to quality estimation that rely on a supervised model.
Movie101: A New Movie Understanding Benchmark (2023.acl-long)

Copied to clipboard

Challenge: Existing methods to narrate movies with no actors are difficult to implement in real situations . a new metric is proposed to provide the best correlation with human evaluation .
Approach: They propose a large-scale Chinese movie benchmark to help visually impaired enjoy movies . they propose metric called Movie Narration Score (MNScore) which achieves best correlation with human evaluation.
Outcome: The proposed method outperforms baselines and the existing methods.
KidsArtBench: Multi-Dimensional Children’s Art Evaluation with Attribute-Aware MLLMs (2026.eacl-long)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) show impressive capabilities across visual–language tasks, but their capacity to evaluate artistic expression remains limited.
Approach: They propose an attribute-specific multi-LoRA approach where each attribute corresponds to a distinct evaluation dimension in the scoring rubric.
Outcome: The proposed approach increases correlation from 0.468 to 0.653 on Qwen2.5-VL-7B, with the largest gains on perceptual dimensions and narrowed gaps on higher-order attributes.
Putting Evaluation in Context: Contextual Embeddings Improve Machine Translation Evaluation (P19-1)

Copied to clipboard

Challenge: Existing evaluation metrics are limited and can be easily portable to new languages.
Approach: They propose a simple unsupervised metric and additional supervised metrics which rely on contextual word embeddings to encode the translation and reference sentences.
Outcome: The proposed model outperforms existing metrics on the WMT 2017 dataset and is more accurate than existing models.
Large Language Models are Not Yet Human-Level Evaluators for Abstractive Summarization (2023.findings-emnlp)

Copied to clipboard

Challenge: ChatGPT and GPT-4 are popular as evaluation metric for complex generative tasks . however, they are not ready as human replacements due to significant limitations .
Approach: They conduct extensive analysis to examine the stability and reliability of LLMs as automatic evaluators for abstractive summarization.
Outcome: The proposed methods outperform the commonly used automatic metrics but are not ready for human evaluation due to significant limitations.
Tokenization and the Noiseless Channel (2023.acl-long)

Copied to clipboard

Challenge: Subword tokenization is a key part of most NLP pipelines, but little is known about why some combinations lead to improved downstream model performance.
Approach: They propose that good tokenizers lead to efficient channel usage . they propose that an optimal encoding assigns extremely long codes to low-frequency subwords .
Outcome: The proposed tokenizers have a very strong correlation with BLEU in machine translation . the proposed function can be used to improve model performance in the downstream task .
KoBE: Knowledge-Based Machine Translation Evaluation (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for machine translation evaluation do not require reference translations.
Approach: They propose a method for machine translation evaluation which does not require reference translations.
Outcome: The proposed method achieves highest correlation with human judgements on 9 out of 18 language pairs from the WMT19 benchmark for evaluation without references.
RCScore: Quantifying Response Consistency in Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Current evaluations of large language models rely on a single instruction template, overlooking models’ sensitivity to instruction style.
Approach: They propose a multi-dimensional framework quantifying how instruction formulation affects model responses by transforming benchmark problems into multiple instruction styles.
Outcome: The proposed framework reveals that instruction style can shift accuracy by 16.7% points.
Do Transformer Models Show Similar Attention Patterns to Task-Specific Human Gaze? (2022.acl-long)

Copied to clipboard

Challenge: We compare attention functions in pre-trained language models to human eye fixation patterns during task-specific reading tasks.
Approach: They compare attention functions in large-scale pre-trained language models to classical cognitive models of human attention by using a dataset with eye-tracking recordings of native speakers of English.
Outcome: The proposed model is as predictive of human eye fixation patterns as classical cognitive models of human attention.
Layer or Representation Space: What Makes BERT-based Evaluation Metrics Robust? (2022.coling-1)

Copied to clipboard

Challenge: Recent embedding-based evaluation metrics for text generation are based on measuring correlation with human evaluations on standard benchmarks.
Approach: They examine the robustness of BERTScore, one of the most popular embedding-based metrics for text generation.
Outcome: The embedding-based metrics that have the highest correlation with human evaluations on a standard benchmark can have the lowest correlation if the amount of input noise or unknown tokens increases.
GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering (2025.coling-main)

Copied to clipboard

Challenge: Existing automated RAG evaluation frameworks overlook important failure modes when using GPT-4 as a judge.
Approach: They propose a novel pipeline to assess the calibration and discrimination capabilities of judge models by using a meta-evaluation benchmark of 144 unit tests to identify key failure modes.
Outcome: The proposed pipeline improves on existing frameworks, while state-of-the-art open-source judges do not generalize to their proposed criteria.
Better Rewards Yield Better Summaries: Learning to Summarise Without References (D19-1)

Copied to clipboard

Challenge: Reinforcement Learning (RL)-based document summarisation systems produce state-of-the-art performance in terms of ROUGE scores, but high summaries receive low human judgement.
Approach: They propose to learn a reward function from human ratings on 2,500 summaries to generate human-appealing summary.
Outcome: The proposed reward function can generate human-appealing summaries without reference summary input.
AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluation methods based on large language models (LLMs) are expensive and lack expertise due to limitations in human expertise.
Approach: They propose an open-source automatic evaluation model with 13B parameters specifically engineered to measure the question-answering proficiency of medical LLMs.
Outcome: The proposed model surpasses baselines in terms of correlation with human judgments.
Adapting Sentence-level Automatic Metrics for Document-level Simplification Evaluation (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies on text simplification have focused on sentence simplification, but these metrics often underperform on longer texts.
Approach: They propose to adapt existing sentence-level metrics for paragraph- or document-level simplification by incorporating a new approach to the evaluation of text simplification metrics.
Outcome: The proposed approach outperforms existing sentence-level metrics in terms of correlation with human judgment and the sensitivity and robustness of various metrics to different types of errors produced by existing systems.
Extracting structure from an LLM - how to improve on surprisal-based models of Human Language Processing (2025.coling-main)

Copied to clipboard

Challenge: Existing computational models capture prediction and reanalysis using Large Language Models (LLMs) and a statistical measure known as ‘surprisal’.
Approach: They propose to extract structural information from Large Language Models and a statistical measure known as ‘surprisal’ to integrate it with their learnt statistics.
Outcome: The proposed model achieved higher correlation with human reading times and better predicted the garden path effect and could distinguish between sentence types with different levels of difficulty.
Competency-Aware Neural Machine Translation: Can Machine Translation Know its Own Translation Quality? (2022.emnlp-main)

Copied to clipboard

Challenge: Neural machine translation models are often criticized for failures that happen without competency awareness.
Approach: They propose a method that extends conventional NMT with a self-estimator to translate a source sentence and estimate its competency.
Outcome: The proposed method performs on translation tasks intact and on quality estimation tasks better than existing methods.
INSTRUCTSCORE: Towards Explainable Text Generation Evaluation with Automatic Feedback (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods to evaluate the quality of language generation do not provide explicit explanation of their verdicts.
Approach: They propose a fine-grained explainable evaluation metric for text generation that harnesses human instruction and implicit knowledge of GPT-4 to fine-tune it.
Outcome: The proposed model outperforms all other unsupervised metrics on translation, captioning, data-to-text, and commonsense generation tasks.
CAMIEval: Enhancing NLG Evaluation through Multidimensional Comparative Instruction-Following Analysis (2025.naacl-long)

Copied to clipboard

Challenge: Evaluating the quality of texts generated by language models has always been a challenging task in natural language processing (NLP).
Approach: They propose a multidimensional comparative evaluation method based on instruction-following that combines relevance, factuality, and adherence with a concrete Chain-of-Thoughts process to enhance the accuracy of evaluations.
Outcome: The proposed method outperforms existing methods in correlation with human evaluations on two NLG evaluation benchmarks.
Facet-Aware Evaluation for Extractive Summarization (2020.acl-main)

Copied to clipboard

Challenge: lexical overlap is a common evaluation metric for extractive summarization, but recent studies reveal its limitations.
Approach: They propose a facet-aware evaluation setup for better assessment of information coverage in extractive summaries.
Outcome: The proposed evaluation setup improves human correlation with extractive summarization datasets and improves comparative analysis.
FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive Summarization (2020.acl-main)

Copied to clipboard

Challenge: Existing automatic metrics do not capture errors in abstractive summarization models.
Approach: They propose an automatic question answering metric for faithfulness that leverages recent advances in reading comprehension.
Outcome: The proposed metric has significantly higher correlation with human faithfulness scores on highly abstracted summaries.
Revisiting Grammatical Error Correction Evaluation and Beyond (2022.emnlp-main)

Copied to clipboard

Challenge: Pretraining-based (PT) evaluation metrics are not effective for training grammatical error correction systems.
Approach: They propose a pretraining-based GEC evaluation metric which only uses PT-based metrics to score the corrected parts of the system.
Outcome: The proposed evaluation metric outperforms existing methods on a CoNLL14 evaluation task.
X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects (2024.naacl-long)

Copied to clipboard

Challenge: X-Eval is a two-stage instruction tuning framework to evaluate text in both seen and unseen aspects customized by end users.
Approach: They introduce a two-stage instruction tuning framework to evaluate text in both seen and unseen aspects customized by end users.
Outcome: The proposed framework improves the model’s ability to follow evaluation instructions and enhances the learning stage to better assess text quality.
HiFi: High-Information Attention Heads Hold for Parameter-Efficient Model Adaptation (2023.acl-long)

Copied to clipboard

Challenge: Existing paradigm to fine-tune parameters of pre-trained language models poses problems in data-scarce and resource-limited scenarios.
Approach: They propose a parameter-efficient fine-tuning method HiFi that fine-tails only the highly informative and strongly correlated attention heads for the specific task.
Outcome: The proposed method obtains state-of-the-art over the prior benchmarks on the GLUE benchmark.
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation (2025.emnlp-main)

Copied to clipboard

Challenge: N-gram-based evaluation metrics are unreliable due to low correlation to human judgments.
Approach: They propose a metric that rewards correct details and penalizes incorrect ones.
Outcome: The proposed metric matches the performance of open-source LLM-based metrics in correlation to human judgments while being far more efficient.
Why Don’t Prompt-Based Fairness Metrics Correlate? (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to assess fairness using prompts have low correlations between fairness metrics.
Approach: They propose a method to enhance the correlation between fairness metrics by using pre-trained language models.
Outcome: The proposed method improves the correlation between fairness metrics by using pre-trained language models.
SANCL: Multimodal Review Helpfulness Prediction with Selective Attention and Natural Contrastive Learning (2022.coling-1)

Copied to clipboard

Challenge: e-commerce has become a research hotspot for review helpfulness prediction . a new approach to help predict helpfulness of multimodal product reviews is proposed .
Approach: They propose a machine learning task to identify helpfulness of multimodal product reviews . they use a probe-based strategy to enforce high attention weights on regions of greater significance .
Outcome: The proposed model achieves state-of-the-art performance with lower memory consumption on two benchmark datasets with three categories.
BLASER: A Text-Free Speech-to-Speech Translation Evaluation Metric (2023.acl-long)

Copied to clipboard

Challenge: End-to-End speech-to speech translation is generally evaluated with text-based metrics . this means generated speech has to be automatically transcribed, making the evaluation dependent on ASR systems.
Approach: They propose a text-free evaluation metric for end-to-end speech-tospeech translation, named BLASER, to avoid the dependency on automatic speech recognition systems.
Outcome: The proposed metric avoids the dependency on automatic speech recognition systems by encoding generated speech segments into a shared embedding space.
Language Model Augmented Relevance Score (2021.acl-long)

Copied to clipboard

Challenge: Existing metrics that compare the candidate with the human reference do not consider the context, resulting in poor correlation with human judgements.
Approach: They propose a language model-aware metric that augments the human reference while considering the context to provide evaluation scores that correlate highly with human judgements.
Outcome: The proposed metric achieves higher correlation with human reference judgements and differentiates well-formed candidates from adversarial samples to a larger degree.
Learning Concept Abstractness Using Weak Supervision (D18-1)

Copied to clipboard

Challenge: Existing methods for inferring abstractness of words and expressions without labeled data are limited and limited.
Approach: They propose a weakly supervised approach for inferring the property of abstractness of words and expressions in the absence of labeled data.
Outcome: The proposed approach obtains high correlation with human labels in the absence of labeled data.
Finding a Balanced Degree of Automation for Summary Evaluation (2021.emnlp-main)

Copied to clipboard

Challenge: Automated summarization metrics are reliable but often poorly correlated with human judgment.
Approach: They propose a semi-automatic to automatic summary evaluation metrics, following the Pyramid human evaluation method.
Outcome: The proposed metrics are semi-automatic to automatic summary evaluation metrics, following the Pyramid human evaluation method.
Semantic Evaluation of Multilingual Data-to-Text Generation via NLI Fine-Tuning: Precision, Recall and F1 scores (2025.findings-acl)

Copied to clipboard

Challenge: KG-to-Text models are prone to errors like Additions and Omissions, and few languages are taken into account since both train and test data are not readily available.
Approach: They propose a multilingual evaluation framework that is reference-less . it allows estimating how much a KG-to-Text Model under- (omission) or over- (addition) generates.
Outcome: The proposed evaluation framework outperforms prior reference-less metrics in correlation with human judgments and provides scores for precision and recall.
SOME: Reference-less Sub-Metrics Optimized for Manual Evaluations of Grammatical Error Correction (2020.coling-main)

Copied to clipboard

Challenge: Existing reference-less metrics are not optimized for manual evaluations of system outputs because no dataset exists for manual analysis.
Approach: They propose a reference-less metric trained on manual evaluations of system outputs for grammatical error correction.
Outcome: The proposed metric improves correlation with manual evaluation in system- and sentence-level meta-evaluation.
Towards Automatic Evaluation of Dialog Systems: A Model-Free Off-Policy Evaluation Approach (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for evaluation of dialog systems are expensive and not scalable . a framework for estimating human evaluation scores is proposed to bridge this gap .
Approach: They propose a framework for estimating human evaluation scores based on off-policy evaluation . they use language quality metrics for single-turn response generation given a fixed context .
Outcome: The proposed framework outperforms existing methods in terms of correlation with human evaluation scores.
CLIPScore: A Reference-free Evaluation Metric for Image Captioning (2021.emnlp-main)

Copied to clipboard

Challenge: Image captioning relies on reference-based automatic evaluations, but references are expensive to collect and comparing against multiple human-authored captions is insufficient.
Approach: They propose a reference-free metric that can be used for automatic caption evaluation without references.
Outcome: The proposed model outperforms existing metrics on image-text compatibility and a reference-augmented version achieves even higher correlation with human judgements.
ACUEval: Fine-grained Hallucination Evaluation and Correction for Abstractive Summarization (2024.findings-acl)

Copied to clipboard

Challenge: Recent-proposed evaluation metrics for large language models have a preference-bias . however, such metrics often lack interpretability and only offer a single score .
Approach: They propose a metric that leverages the power of large language models to perform two sub-tasks: decomposing summaries into atomic content units and validating them against the source document.
Outcome: The proposed metric improves faithfulness scores on three summarization evaluation benchmarks by 3% compared to the next-best metric.
imapScore: Medical Fact Evaluation Made Easy (2024.findings-acl)

Copied to clipboard

Challenge: Automated evaluation of natural language generation tasks fails to focus on medical QA because of the diversity in medical terminology.
Approach: They propose a new data structure, imap, to capture key information in questions and answers.
Outcome: The proposed model outperforms state-of-the-art metrics in correlation with human scores.
Q2: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering (2021.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation methods for factual consistency in knowledge-grounded dialogues are unreliable and limit their applicability.
Approach: They propose an automatic evaluation metric for factual consistency in knowledge-grounded dialogue using automatic question generation and question answering.
Outcome: The proposed evaluation metric consistently shows higher correlation with human judgements.
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for reward models show a weak correlation with performance of optimized policies . existing benchmarks do not accurately assess the true capabilities of reward models .
Approach: They explore how reward overoptimization captures how well a reward model aligns with human preferences and the dynamics of the learning signal it provides to the policy.
Outcome: The proposed benchmarks show that reward overoptimization is a weak factor . the high correlation with degree of overoptimalization leads to lower correlation with downstream performance .
ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for evaluating factual consistency are primarily designed for short summaries of isolated code snippets.
Approach: They propose a reference-free and fine-grained method for evaluating factual consistency in real-world code summaries.
Outcome: The proposed method achieves highest correlation with human judgment among 13 baselines, improving 15-18% over the previous state-of-the-art.
VIFIDEL: Evaluating the Visual Fidelity of Image Descriptions (P19-1)

Copied to clipboard

Challenge: Existing methods for evaluating image description generation systems are subjective and expensive to scale.
Approach: They propose a new image-aware metric for evaluating image description generation systems . it estimates the faithfulness of a generated caption with respect to the content of the actual image .
Outcome: The proposed metric achieves high correlation with human judgments on two well-known datasets and is competitive with metrics that depend on and rely exclusively on human references.
FactSpotter: Evaluating the Factual Faithfulness of Graph-to-Text Generation (2023.findings-emnlp)

Copied to clipboard

Challenge: Graph-to-text (G2T) generation is an important task in natural language generation as it renders graphs accessible to non-technical users in downstream applications such as question answering.
Approach: They propose a metric that correctly identifies factual faithfulness and uses it to determine if a triple is present in a generated text.
Outcome: The proposed metric achieves highest correlation with human annotations on data correctness, data coverage, and relevance.
Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs (2024.findings-emnlp)

Copied to clipboard

Challenge: Current evaluation practices of open domain dialogue systems are still highly dependent on human evaluation.
Approach: They propose to use an annotated dataset to evaluate chatbots using large language models.
Outcome: The proposed model improves over few-shot inferences on a GPT-3.5 generated dialogue dataset.
Mediocrity is the key for LLM as a Judge Anchor Selection (2026.acl-long)

Copied to clipboard

Challenge: a poor selection of an anchor can dramatically reduce correlation with human rankings . traditional reference-based metrics are often ill-suited for open-ended generation .
Approach: They evaluate 22 different anchors on a Arena-Hard-v2.0 dataset and quantify the effect size of anchor selection.
Outcome: The proposed model is better or worse than all other models, but it is rarely indicative of the relative ranking of the models.
BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics (2023.acl-long)

Copied to clipboard

Challenge: Existing benchmarks measure the correlation with human judgements of faithfulness on model-generated summaries, but they are insufficient for diagnosing whether metrics are consistent, effective on human-written texts, and sensitive to different error types.
Approach: They propose to use unfaithful minimal pairs to measure the consistency of automatic faithfulness metrics by comparing human-written summary pairs with a dataset of 889 human-writing, minimally different summary pairs.
Outcome: The proposed benchmarks show that the most discriminative metrics tend not to be the most consistent, and that the best performing metrics are sensitive to errors.
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era (2025.findings-acl)

Copied to clipboard

Challenge: Image captioning has been a challenge for vision-language researchers for decades . current VLMs focus on tasks like visual question answering (YA) but image captioning is not as advanced as expected.
Approach: They evaluate VLMs' performance on image captioning using human annotations . they find that some metrics show high caption-level agreement with humans .
Outcome: The proposed model outperforms open-source models on image captioning . it achieves 93.4% correlation with human rankings at $4 per test .
X-ACE: Explainable and Multi-factor Audio Captioning Evaluation (2024.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics for automated audio captioning only provide an overall score . current evaluation checklists are inadequate to characterize the nuanced differences .
Approach: They propose an explainable and multi-factor audio captioning evaluation paradigm . they define sound event, source, attribute and relation as four factors tailored for the audio description .
Outcome: The proposed evaluation paradigm improves the quality of audio captions . it can detect mismatches and align with human perception, the authors show .
Simplicity Level Estimate (SLE): A Learned Reference-Less Metric for Sentence Simplification (2023.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics conflate simplicity with correlated attributes such as fluency or meaning preservation.
Approach: They propose a new learning evaluation metric that focuses on simplicity outperforming most existing metrics in terms of correlation with human judgements.
Outcome: The proposed metric outperforms most existing metrics in terms of correlation with human judgements.
Guilt by Association: Emotion Intensities in Lexical Representations (2021.emnlp-main)

Copied to clipboard

Challenge: linguistic models have a higher correlation with human ground truth ratings than labeled data . word vectors have often been evaluated on standard word relatedness benchmarks .
Approach: They propose to use unsupervised, supervised, and finally supervised methods to extract emotional associations from pretrained vectors and models.
Outcome: The proposed method shows higher correlation with ground truth ratings than state-of-the-art lexicons based on labeled data.
Can LLMs Understand the Implication of Emphasized Sentences in Dialogue? (2024.findings-emnlp)

Copied to clipboard

Challenge: Emphasis is a crucial component in human communication, which indicates speaker’s intention and implication beyond pure text in dialogue.
Approach: They propose a benchmark dataset with annotated dialogue samples capturing the implications of emphasis.
Outcome: The proposed evaluation pipeline achieves high correlation with human scoring and commercial LLMs perform better than open-source LLM.
Digging Errors in NMT: Evaluating and Understanding Model Errors from Partial Hypothesis Space (2022.emnlp-main)

Copied to clipboard

Challenge: Current evaluation of neural machine translation systems is limited by one best hypothesis and search errors brought by heuristic decoding algorithms.
Approach: They propose a new evaluation protocol which defines model errors with model’s ranking capability over hypothesis space and Monte Carlo sampling evaluation to tackle the problem of exponentially large space.
Outcome: The proposed evaluation protocol is consistent with what is currently used in the field and is consistent to what is being proposed.
RISE: Leveraging Retrieval Techniques for Summarization Evaluation (2023.findings-acl)

Copied to clipboard

Challenge: Summarization evaluation approaches have relied on ROUGE for summarization, but they fall short of human evaluations.
Approach: They propose a new approach to evaluate summaries by leveraging retrieval techniques . they use a dual-encoder retrieval setup to train a retrieval task .
Outcome: The proposed method outperforms existing methods on two document summarization benchmarks and a long document summmarization test.
SLIDE: A Framework Integrating Small and Large Language Models for Open-Domain Dialogues Evaluation (2024.findings-acl)

Copied to clipboard

Challenge: Existing approaches to evaluate open domain dialogues have a one-to-many problem . existing approaches lack commonsense reasoning biases and perform poorly in domain-specific scenarios.
Approach: They propose a framework that leverages both a small, specialised model and LLMs for the evaluation of open-domain dialogues.
Outcome: The proposed framework achieves state-of-the-art performance in both classification and evaluation tasks and exhibits better correlation with human judgements.
Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization (2025.acl-long)

Copied to clipboard

Challenge: n-gram-based metrics are considered indicative (even if imperfect) of human evaluation for English, but their suitability for other languages remains unclear.
Approach: They systematically assess evaluation metrics for generation for languages and tasks using n-gram-based and neural-based metrics.
Outcome: The proposed evaluation suite is based on eight languages from four typological families and shows that it is sensitivity to the language type at hand.
The Progress Illusion: Revisiting meta-evaluation standards of LLM evaluators (2025.findings-emnlp)

Copied to clipboard

Challenge: LLM judges have gained popularity as an inexpensive and performant substitute for human evaluation.
Approach: They revisit meta-evaluations of LLM evaluators under a setting that more closely aligns with practice by examining evaluers’ ability to distinguish test system pairs that are closer in capability.
Outcome: The proposed meta-evaluation setting is significantly different from the use of human evaluations.
The Linearity of the Effect of Surprisal on Reading Times across Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: a large amount of insight into human language processing can be gleaned by studying word-by-word processing difficulty.
Approach: They extend the study by examining eyetracking corpora of seven languages . they find evidence for superlinearity in some languages, but highly sensitive to language models .
Outcome: The study extends existing studies on english to Danish, Dutch, English, German, Japanese, Mandarin, and Russian.
Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that MT metrics return assessments as scalar scores that are difficult to interpret, posing a challenge to making informed design choices.
Approach: They propose an interpretable evaluation framework that evaluates MT metrics in two scenarios that serve as proxies for filtering and translation re-ranking use cases.
Outcome: The proposed framework offers clearer insights than correlation with human judgments.
PWESuite: Phonetic Word Embeddings and Tasks They Facilitate (2024.lrec-main)

Copied to clipboard

Challenge: Existing word embedding methods overlook phonetic information that is crucial for many tasks.
Approach: They propose three methods that use articulatory features to build phonetically informed word embeddings.
Outcome: The proposed methods improve word retrieval and correlation with sound similarity and on rhyme and cognate detection tasks.
ReadLet: A Dataset for Oral, Visual and Tactile Text Reading Data of Early and Mature Readers (2024.lrec-main)

Copied to clipboard

Challenge: The paper presents the design and construction of a time-aligned multimodal dataset for reading research, including multiple time-aligned temporal signals elicited with four experimental trials of connected text reading by both child and adult readers.
Approach: They propose to use a time-stamped multimodal dataset to analyze time-aligned temporal signals elicited by connected text reading by both child and adult readers.
Outcome: The proposed dataset includes multiple time-aligned temporal signals elicited with four experimental trials of connected text reading by both child and adult readers.
Agent-as-Judge for Factual Summarization of Long Narratives (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated near-human performance in summarization tasks based on traditional metrics such as ROUGE and BERTScore . however, these metrics do not adequately capture critical aspects of summarizing quality, such as factual accuracy, especially for long narratives.
Approach: They propose a framework that evaluates and refines factuality in narrative summarization by leveraging a Character Knowledge Graph extracted from input narrative.
Outcome: The proposed framework evaluates factuality and provides actionable guidance for refinement.
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies show that LLM-based agents struggle to perform in zero-shot scenarios.
Approach: They propose a framework to quantify the behavior gap between AI agents and human experts . they propose to examine discrepancies in dialog acts, tool usage, and knowledge utilization .
Outcome: The proposed framework measures the behavior gap between AI agents and human experts on task-oriented dialogs.
xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation Metrics (2024.emnlp-main)

Copied to clipboard

Challenge: State-of-the-art trainable machine translation evaluation metrics rely on large encoders . this makes them computationally expensive and inaccessible to researchers with limited resources.
Approach: They propose a method to extract knowledge stored in large encoders and a pipeline for efficient black-box distillation.
Outcome: The proposed model surpasses COMET-22 and BLEURT-20 on the WMT22 dataset by 6.4%.
IMPARA-GED: Grammatical Error Detection is Boosting Reference-free Grammatical Error Quality Estimator (2025.findings-acl)

Copied to clipboard

Challenge: Existing reference-free automatic grammatical error correction methods do not correlate with human evaluation.
Approach: They propose a reference-free automatic grammatical error correction evaluation method with enhanced gramma-ed capabilities.
Outcome: The proposed method achieves highest correlation with human evaluations on a meta-evaluation dataset.
InspireDebate: Multi-Dimensional Subjective-Objective Evaluation-Guided Reasoning and Optimization for Debating (2025.acl-long)

Copied to clipboard

Challenge: Existing LLMs focus on responding to specific arguments while neglecting objective assessments such as authenticity and logical validity.
Approach: They propose a multi-dimensional evaluation system and an optimized debating framework . they propose to use coT reasoning enhancement, web-based Retrieval Augmented Generation to optimize across various dimensions.
Outcome: The proposed framework outperforms baseline models in argument quality assessment and debate process simulation by 57%.
QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing video-to-text summarization evaluation methods depend heavily on human-written reference summaries.
Approach: They propose a reference-free metric evaluating candidate summaries directly against source videos through multimodal question answering.
Outcome: The proposed metric assesses candidate summaries directly against source videos through multimodal question answering.
Reliability Crisis of Reference-free Metrics for Grammatical Error Correction (2025.findings-emnlp)

Copied to clipboard

Challenge: Reference-free evaluation metrics for grammatical error correction have high correlation with human judgments, but they are not designed to evaluate adversarial systems that aim to obtain unjustifiably high scores.
Approach: They propose adversarial attack strategies for four reference-free metrics . they propose SOME, Scribendi, IMPARA, and LLM-based metrics based on these metrics a .
Outcome: The proposed attacks outperform the current state-of-the-art for four reference-free metrics .
LiTransProQA: An LLM-based Literary Translation Evaluation Metric with Professional Question Answering (2025.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics for literature prioritize mechanical accuracy over artistic expression . this bias could result in an irreversible decline in translation quality and cultural authenticity .
Approach: They propose a novel, reference-free, LLM-based question-answering framework for literary translation evaluation.
Outcome: a novel, reference-free, LLM-based question-answering framework is developed for literary translation evaluation.
VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis (2026.acl-long)

Copied to clipboard

Challenge: Existing metrics for caption evaluation lack factual accuracy and limited context handling . VC-Inspector provides reproducible, fact-aware alternative that aligns closely with human judgments.
Approach: They propose a lightweight, open-source large multimodal model for reference-free evaluation of video captions with a focus on factual accuracy.
Outcome: Experiments show that VC-Inspector can generalize across diverse domains and improve on existing metrics.
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces (2025.acl-long)

Copied to clipboard

Challenge: Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban aerial spaces remain to be explored.
Approach: They propose a benchmark to evaluate whether large multimodal models can process continuous first-person visual observations like humans.
Outcome: The proposed model can process first-person visual observations like humans, enabling recall, perception, reasoning, and navigation.
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown promise in generating visualizations from natural language, but lack of comprehensive benchmarks limits their capabilities.
Approach: They propose a framework that jointly refines the textual answer and visualization code to improve GPT-4o's pass rate from 26% to 42% over direct approach.
Outcome: The proposed framework increases GPT-4o’s pass rate from 26% to 42% over the direct approach and improves chart quality.
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models (2026.acl-long)

Copied to clipboard

Challenge: closed-ended question-based benchmarks struggle with saturation as newer models emerge . crowd-sourced leaderboards rely on costly and slow human judges .
Approach: They propose a framework that leverages collective intelligence from all large language models to evaluate each other.
Outcome: a new framework enables a democratic, pairwise evaluation of all large language models . it achieves 97% correlation with human judgements, while significantly reducing the cost.
TounsiBench: Benchmarking Large Language Models for Tunisian Arabic (2025.emnlp-main)

Copied to clipboard

Challenge: a dataset of Tunisian Arabic instructions and prompts is used to evaluate LLMs' ability to understand and generate responses in Tunisia . we assess the quality, correctness, relevance, and dialectal adherence of LLM responses .
Approach: They propose a benchmark for evaluating the capabilities of large language models in Tunisian Arabic . they use a dataset of Tunisia Arabic instructions and prompts to evaluate their models .
Outcome: The proposed model can judge quality, correctness, relevance, and dialectal adherence . the model can also generate a leaderboard for the Tunisian Arabic language .
Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment (2026.acl-long)

Copied to clipboard

Challenge: Existing LAM benchmarks with thousands of examples create substantial computational barriers.
Approach: They examine whether subsets can reliably evaluate large audio models . they find that subset of 50 examples can achieve over 0.93 Pearson correlation with full benchmark .
Outcome: The proposed method outperforms the full benchmark and subset selection methods.
DualFact+: A Multimodal Fact Verification Framework for Procedural Video Captioning (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics fail to evaluate factual correctness in procedural video captions . Existing metrics rely on lexical overlap or holistic semantic similarity, but miss role-specific omissions resulting in hallucinations .
Approach: They propose a role-aware, fact-level evaluation framework that distinguishes conceptual facts from contextual facts.
Outcome: Experiments show that state-of-the-art captioning models produce fluent but incomplete descriptions with systematic errors.
When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection (2026.acl-long)

Copied to clipboard

Challenge: Personalized MGT detection remains largely underexplored due to personalization challenges . large language models (LLMs) can imitate personal writing styles, but they can generate fake news and misinformation.
Approach: They propose a benchmark to evaluate detector robustness under personalization . they attribute this limitation to a feature-inversion trap that flips the effect in personalized contexts .
Outcome: The proposed framework predicts detector robustness under personalization with an 85% correlation to actual results.
TabReX: Tabular Referenceless eXplainable Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing metrics for evaluating the quality of tables generated by large language models flatten tables into text, ignoring structure or relying on fixed references that limit generalization.
Approach: They propose a reference-less framework for evaluating tabular generation via graph-based reasoning . tabReX converts source text and generated tables into canonical knowledge graphs .
Outcome: The proposed framework provides a high correlation with expert rankings and stable under harder perturbations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations