Challenge: Existing studies on topic models lack a correlation between automated coherence metrics and human judgement.
Approach: They propose a sampling approach to mine topics for metric evaluation and extend the analysis to measure topical differences between corpora.
Outcome: The proposed method extends to measure topical differences between corpora and human judgement by using extensive user study.

Similar Papers

How to Find Strong Summary Coherence Measures? A Toolbox and a Comparative Study for Summary Coherence Measure Evaluation (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to evaluate summary coherence are often evaluated using disparate datasets and metrics.
Approach: They propose to use automatic evaluation to evaluate coherence of summaries by selecting high-scoring candidates.
Outcome: The proposed methods show that they can perform better on an even playing field.
Re-Examining System-Level Correlations of Automatic Summarization Evaluation Metrics (2022.naacl-main)

Copied to clipboard

Challenge: Existing definitions of system-level correlations are inconsistent with how they are used to evaluate systems.
Approach: They propose to calculate correlations only on pairs of systems separated by small differences in automatic scores . they propose to use the full test set instead of the subset of summaries judged by humans .
Outcome: The proposed changes improve the accuracy of the estimated correlations on pairs of systems separated by small differences in automatic scores.
Characterizing the Confidence of Large Language Model-Based Automatic Evaluation Metrics (2024.eacl-short)

Copied to clipboard

Challenge: Recent studies have focused on using Large Language Models (LLMs) to evaluate NLP tasks automatically.
Approach: They characterize LLM evaluators’ confidence in ranking candidate NLP models and develop a configurable Monte Carlo simulation method to compensate for loss of correlation.
Outcome: The proposed method can reach 95% confidence rankings of candidate models with reasonable evaluation set sizes.
Contextualized Topic Coherence Metrics (2024.findings-eacl)

Copied to clipboard

Challenge: Existing topic models that estimate the interpretability of topics are difficult to compare due to their nature as unsupervised models.
Approach: They propose to use contextualized topic coherence metrics to simulate human-centered coherency evaluation while maintaining the efficiency of other automated methods.
Outcome: The proposed metrics better reflect human judgment on topics extracted from short text collections by avoiding highly scored topics that are meaningless to humans.
A Measure of the System Dependence of Automated Metrics (2025.acl-short)

Copied to clipboard

Challenge: Recent advances in machine translation evaluations are expensive and time-intensive.
Approach: They propose a method to evaluate the correlation between human and metric scores . they argue that it is equally important to ensure that metrics treat all systems fairly and consistently.
Outcome: The proposed method ignores a central requirement of the evaluation process, and ignores the need for a thorough evaluation procedure.
Evaluation of Thematic Coherence in Microblogs (2021.acl-long)

Copied to clipboard

Challenge: Recent work on grouping together views about tweets expressing opinions about the same entities has been criticized for their lack of thematic coherence.
Approach: They propose to use a corpus of microblogs representing opinions about the same topics within the same time window to evaluate thematic coherence.
Outcome: The proposed method outperforms surface level metrics, topic model coherence and text generation metrics (TGMs) but is not as reliable as TGMs due to being less sensitive to time windows.
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies have not investigated the differences between different correlation measures in meta-evaluation.
Approach: They analyze 12 common correlation measures using real-world data from six widely-used NLG evaluation datasets and 32 evaluation metrics.
Outcome: The proposed measures exhibit the best performance in discriminative power and ranking consistency . the measures using system-level grouping or Kendall correlation are the least sensitive to score granularity .
Automatic Evaluation of Local Topic Quality (P19-1)

Copied to clipboard

Challenge: Topic models are evaluated with global topic distributions but without local topic assignments.
Approach: They propose a task to elicit human judgments of token-level topic assignments . they propose to use global metrics to evaluate topic models at a local level .
Outcome: The proposed task elicits human judgments of token-level topic assignments . global metrics agree poorly with human assignments, the authors show .
Studying Summarization Evaluation Metrics in the Appropriate Scoring Range (P19-1)

Copied to clipboard

Challenge: Existing evaluation metrics are compared based on their ability to correlate with humans, but they disagree in the higher-scoring range in which current systems operate.
Approach: They show that evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate.
Outcome: The evaluation metrics which behave similarly on these datasets strongly disagree in the higher-scoring range in which current systems operate.
GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue Systems (2020.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics only consider surface features or utterance-level semantics, without explicitly considering the fine-grained topic transition dynamics of dialogue flows.
Approach: They propose a graph-enhanced evaluation metric GRADE to evaluate dialogue coherence . GRADE incorporates utterance-level contextualized representations and fine-grained topic-level graph representations to improve communication logic.
Outcome: The proposed evaluation metric outperforms state-of-the-art metrics on measuring diverse dialogue models in terms of Pearson and Spearman correlations with human judgments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations