Large-Scale Correlation Analysis of Automated Metrics for Topic Models (2023.acl-long)
Copied to clipboard
| Challenge: | Existing studies on topic models lack a correlation between automated coherence metrics and human judgement. |
| Approach: | They propose a sampling approach to mine topics for metric evaluation and extend the analysis to measure topical differences between corpora. |
| Outcome: | The proposed method extends to measure topical differences between corpora and human judgement by using extensive user study. |
Similar Papers
How to Find Strong Summary Coherence Measures? A Toolbox and a Comparative Study for Summary Coherence Measure Evaluation (2022.coling-1)
Copied to clipboard
| Challenge: | Existing methods to evaluate summary coherence are often evaluated using disparate datasets and metrics. |
| Approach: | They propose to use automatic evaluation to evaluate coherence of summaries by selecting high-scoring candidates. |
| Outcome: | The proposed methods show that they can perform better on an even playing field. |
Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing definitions of system-level correlations are inconsistent with how they are used to evaluate systems. |
| Approach: | They propose to calculate correlations only on pairs of systems separated by small differences in automatic scores . they propose to use the full test set instead of the subset of summaries judged by humans . |
| Outcome: | The proposed changes improve the accuracy of the estimated correlations on pairs of systems separated by small differences in automatic scores. |
Characterizing the Confidence of Large Language Model-Based Automatic Evaluation Metrics (2024.eacl-short)
Copied to clipboard
| Challenge: | Recent studies have focused on using Large Language Models (LLMs) to evaluate NLP tasks automatically. |
| Approach: | They characterize LLM evaluators’ confidence in ranking candidate NLP models and develop a configurable Monte Carlo simulation method to compensate for loss of correlation. |
| Outcome: | The proposed method can reach 95% confidence rankings of candidate models with reasonable evaluation set sizes. |
Contextualized Topic Coherence Metrics (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing topic models that estimate the interpretability of topics are difficult to compare due to their nature as unsupervised models. |
| Approach: | They propose to use contextualized topic coherence metrics to simulate human-centered coherency evaluation while maintaining the efficiency of other automated methods. |
| Outcome: | The proposed metrics better reflect human judgment on topics extracted from short text collections by avoiding highly scored topics that are meaningless to humans. |
A Measure of the System Dependence of Automated Metrics (2025.acl-short)
Copied to clipboard
| Challenge: | Recent advances in machine translation evaluations are expensive and time-intensive. |
| Approach: | They propose a method to evaluate the correlation between human and metric scores . they argue that it is equally important to ensure that metrics treat all systems fairly and consistently. |
| Outcome: | The proposed method ignores a central requirement of the evaluation process, and ignores the need for a thorough evaluation procedure. |
Evaluation of Thematic Coherence in Microblogs (2021.acl-long)
Copied to clipboard
| Challenge: | Recent work on grouping together views about tweets expressing opinions about the same entities has been criticized for their lack of thematic coherence. |
| Approach: | They propose to use a corpus of microblogs representing opinions about the same topics within the same time window to evaluate thematic coherence. |
| Outcome: | The proposed method outperforms surface level metrics, topic model coherence and text generation metrics (TGMs) but is not as reliable as TGMs due to being less sensitive to time windows. |
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing studies have not investigated the differences between different correlation measures in meta-evaluation. |
| Approach: | They analyze 12 common correlation measures using real-world data from six widely-used NLG evaluation datasets and 32 evaluation metrics. |
| Outcome: | The proposed measures exhibit the best performance in discriminative power and ranking consistency . the measures using system-level grouping or Kendall correlation are the least sensitive to score granularity . |
Automatic Evaluation of Local Topic Quality (P19-1)
Copied to clipboard
Jeffrey Lund, Piper Armstrong, Wilson Fearn, Stephen Cowley, Courtni Byun, Jordan Boyd-Graber, Kevin Seppi
| Challenge: | Topic models are evaluated with global topic distributions but without local topic assignments. |
| Approach: | They propose a task to elicit human judgments of token-level topic assignments . they propose to use global metrics to evaluate topic models at a local level . |
| Outcome: | The proposed task elicits human judgments of token-level topic assignments . global metrics agree poorly with human assignments, the authors show . |
Studying Summarization Evaluation Metrics in the Appropriate Scoring Range (P19-1)
Copied to clipboard
| Challenge: | Existing evaluation metrics are compared based on their ability to correlate with humans, but they disagree in the higher-scoring range in which current systems operate. |
| Approach: | They show that evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate. |
| Outcome: | The evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate. |
GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue Systems (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation metrics only consider surface features or utterance-level semantics, without explicitly considering the fine-grained topic transition dynamics of dialogue flows. |
| Approach: | They propose a graph-enhanced evaluation metric GRADE to evaluate dialogue coherence . GRADE incorporates utterance-level contextualized representations and fine-grained topic-level graph representations to improve communication logic. |
| Outcome: | The proposed evaluation metric outperforms state-of-the-art metrics on measuring diverse dialogue models in terms of Pearson and Spearman correlations with human judgments. |