Jeffrey Lund, Piper Armstrong, Wilson Fearn, Stephen Cowley, Courtni Byun, Jordan Boyd-Graber, Kevin Seppi
| Challenge: | Topic models are evaluated with global topic distributions but without local topic assignments. |
| Approach: | They propose a task to elicit human judgments of token-level topic assignments . they propose to use global metrics to evaluate topic models at a local level . |
| Outcome: | The proposed task elicits human judgments of token-level topic assignments . global metrics agree poorly with human assignments, the authors show . |
Similar Papers
Improving the TENOR of Labeling: Re-evaluating Topic Models for Content Analysis (2024.eacl-long)
Copied to clipboard
Zongxia Li, Andrew Mao, Daniel Stephens, Pranav Goel, Emily Walpole, Alden Dima, Juan Fung, Jordan Boyd-Graber
| Challenge: | Existing evaluation metrics such as coherence and coherency are inadequate for neural topic models. |
| Approach: | They conduct the first evaluation of neural, supervised and classical topic models in an interactive task-based setting. |
| Outcome: | The proposed model performs better on cluster evaluation metrics and human evaluations than classical models on real-world tasks. |
Contextualized Topic Coherence Metrics (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing topic models that estimate the interpretability of topics are difficult to compare due to their nature as unsupervised models. |
| Approach: | They propose to use contextualized topic coherence metrics to simulate human-centered coherency evaluation while maintaining the efficiency of other automated methods. |
| Outcome: | The proposed metrics better reflect human judgment on topics extracted from short text collections by avoiding highly scored topics that are meaningless to humans. |
Large-Scale Correlation Analysis of Automated Metrics for Topic Models (2023.acl-long)
Copied to clipboard
| Challenge: | Existing studies on topic models lack a correlation between automated coherence metrics and human judgement. |
| Approach: | They propose a sampling approach to mine topics for metric evaluation and extend the analysis to measure topical differences between corpora. |
| Outcome: | The proposed method extends to measure topical differences between corpora and human judgement by using extensive user study. |
Are Neural Topic Models Broken? (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluation paradigms are often divorced from real-world use . recent results have challenged the validity of the prevailing model evaluation paradigm . |
| Approach: | They show that neural topic models fare worse in both respects compared to an established classical method. |
| Outcome: | The proposed method outperforms the members of the ensemble in both respects. |
Topic Model or Topic Twaddle? Re-evaluating Semantic Interpretability Measures (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing methods for topic model evaluation use automated measures modeled on human evaluation tests that are dissimilar to applied usage. |
| Approach: | They propose to use a novel experimental framework to evaluate topic models and assess their coherence for specialized collections in an applied setting. |
| Outcome: | The proposed framework is reflective of human evaluations using open labeling, typical of applied research. |
Topic Intrusion for Automatic Topic Model Evaluation (D18-1)
Copied to clipboard
| Challenge: | Topic coherence is increasingly being used to evaluate topic models and filter topics for end-user applications. |
| Approach: | They propose to use topic intrusion to guess an outlier topic given a document and a few topics to automate the task. |
| Outcome: | The proposed method improves upon the state-of-the-art method and shows it can be used as an alternative to topic perplexity evaluation. |
Revisiting Automated Topic Model Evaluation with Large Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Topic models are an unsupervised dimensionality reduction technique that help organize large text collections. |
| Approach: | They propose to use large language models to evaluate document output and determine optimal number of topics. |
| Outcome: | The proposed model performs better on coherence ratings of word sets than on intrustion detection. |
ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering (2025.acl-long)
Copied to clipboard
| Challenge: | Topic models and document clustering evaluations often use automated metrics that align poorly with human preferences or require expert labels that are intractable to scale. |
| Approach: | They propose a protocol for evaluating topic models and document clustering evaluations that uses crowdworker annotations to validate automated proxies. |
| Outcome: | The proposed protocol is scalable and easy to adapt to an LLM prompt. |
Evaluating Dynamic Topic Models (2024.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation measures to evaluate the progression of topics in dynamic topic models (DTMs) are difficult due to their unsupervised nature, but are crucial for detecting trends in time-indexed documents. |
| Approach: | They propose to combine topic quality and temporal consistency to evaluate the progression of topics over time in dynamic topic models. |
| Outcome: | The proposed measure correlates well with human judgment and can be used to identify changing topics and evaluate different models and LLMs. |
Evaluation of Thematic Coherence in Microblogs (2021.acl-long)
Copied to clipboard
| Challenge: | Recent work on grouping together views about tweets expressing opinions about the same entities has been criticized for their lack of thematic coherence. |
| Approach: | They propose to use a corpus of microblogs representing opinions about the same topics within the same time window to evaluate thematic coherence. |
| Outcome: | The proposed method outperforms surface level metrics, topic model coherence and text generation metrics (TGMs) but is not as reliable as TGMs due to being less sensitive to time windows. |