| Challenge: | a recent study shows that topic models that highlight differences in authors are often not accurate . authors show that subsampling words that are highly correlated with metadata can reduce topic-metadata correlation . |
| Approach: | They propose three metrics for identifying topics that are highly correlated with metadata . they find that subsampling words causes topic-metadata correlation, improve topic stability . authors propose to use topic models to infer word distributions that correspond to recognizable themes . |
| Outcome: | The proposed model can predict which words cause the phenomenon and improve topic stability and quality. |
Similar Papers
Practical Correlated Topic Modeling and Analysis via the Rectified Anchor Word Algorithm (D19-1)
Copied to clipboard
| Challenge: | spectral topic models lack reliability in real data and lack of practical implementations. |
| Approach: | They propose to use a spectral topic inference method to infer correlations between topics in real data and a matrix-based approach to inference. |
| Outcome: | The proposed method outperforms tensor-based methods and probabilistic methods in real data and provides a complete guide to correlated topic modeling. |
Neural Models for Documents with Metadata (P18-1)
Copied to clipboard
| Challenge: | specialized models are often used to model text corpora without metadata . specialized algorithms are not widely used in the digital humanities and political science fields . |
| Approach: | They propose a general neural framework based on topic models to enable customization of metadata. |
| Outcome: | The proposed framework achieves strong performance with a manageable tradeoff between perplexity, coherence, and sparsity. |
Are Neural Topic Models Broken? (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluation paradigms are often divorced from real-world use . recent results have challenged the validity of the prevailing model evaluation paradigm . |
| Approach: | They show that neural topic models fare worse in both respects compared to an established classical method. |
| Outcome: | The proposed method outperforms the members of the ensemble in both respects. |
Improving Entity Linking by Modeling Latent Relations between Mentions (P18-1)
Copied to clipboard
| Challenge: | Entity linking systems often exploit relations between textual mentions to decide if the linking decisions are compatible. |
| Approach: | They treat relations as latent variables while optimizing the neural entity-linking model without supervision. |
| Outcome: | The proposed model outperforms its relation-agnostic version and significantly outperformed its relational version. |
Inflating Topic Relevance with Ideology: A Case Study of Political Ideology Bias in Social Topic Detection Models (2020.coling-main)
Copied to clipboard
| Challenge: | a study examines the impact of political ideology biases in training data . topic detection methods may contain or propagate certain biase resulting in a skewed data collection . |
| Approach: | They propose to learn a text representation that is invariant to political ideology while still judging topic relevance. |
| Outcome: | The proposed model can be invariant to political ideology while still judging topic relevance. |
Stubborn Lexical Bias in Data and Models (2023.findings-acl)
Copied to clipboard
| Challenge: | Recent work has focused on spurious correlations between features and labels in training data . but, we find strong evidence of corresponding bias in the trained models . |
| Approach: | They propose a method to reduce spurious correlations in training data by reweighting it using a large pool of extracted features. |
| Outcome: | The proposed method reduces spurious correlations in training data, but still finds strong evidence of bias in trained models. |
Boosting Entity Linking Performance by Leveraging Unlabeled Documents (P19-1)
Copied to clipboard
| Challenge: | a new approach to entity linking relies on unlabeled documents and Wikipedia . a supervised approach uses only natural information, such as unlabed documents . |
| Approach: | They propose a method which exploits only naturally occurring information . they construct a high recall list of candidate entities for each mention in an unlabeled document . |
| Outcome: | The proposed model outperforms fully-supervised state-of-the-art systems on standard test sets. |
We Need to Measure Data Diversity in NLP — Better and Broader (2025.emnlp-main)
Copied to clipboard
| Challenge: | Language models exhibit remarkable natural language understanding and generation capabilities, but they have serious flaws, such as societal biases and spurious correlations. |
| Approach: | They argue that interdisciplinary perspectives are essential for developing more fine-grained and valid measures of data diversity. |
| Outcome: | The proposed measures are based on interdisciplinary perspectives and include a variety of datasets. |
PhraseCTM: Correlated Topic Modeling on Phrases within Markov Random Fields (P18-2)
Copied to clipboard
| Challenge: | Recent phrase-level topic models are unable to capture the correlation structure among the discovered topics. |
| Approach: | They propose a phrase-level topic model PhraseCTM and a method to find out the correlations of topics at phrase level. |
| Outcome: | The proposed method shows that correlated topic modeling is a good way to interpret themes of corpus. |
Tethering Broken Themes: Aligning Neural Topic Models with Labels and Authors (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent studies suggest that topic models do not align well with human intentions. |
| Approach: | They propose a method to align neural topic models with both labels and authorship information. |
| Outcome: | The proposed method improves existing models in terms of topic quality and alignment. |