Authorless Topic Models: Biasing Models Away from Known Structure (C18-1)

Copied to clipboard

Challenge: a recent study shows that topic models that highlight differences in authors are often not accurate . authors show that subsampling words that are highly correlated with metadata can reduce topic-metadata correlation .
Approach: They propose three metrics for identifying topics that are highly correlated with metadata . they find that subsampling words causes topic-metadata correlation, improve topic stability . authors propose to use topic models to infer word distributions that correspond to recognizable themes .
Outcome: The proposed model can predict which words cause the phenomenon and improve topic stability and quality.

Similar Papers

Practical Correlated Topic Modeling and Analysis via the Rectified Anchor Word Algorithm (D19-1)

Copied to clipboard

Challenge: spectral topic models lack reliability in real data and lack of practical implementations.
Approach: They propose to use a spectral topic inference method to infer correlations between topics in real data and a matrix-based approach to inference.
Outcome: The proposed method outperforms tensor-based methods and probabilistic methods in real data and provides a complete guide to correlated topic modeling.
Neural Models for Documents with Metadata (P18-1)

Copied to clipboard

Challenge: specialized models are often used to model text corpora without metadata . specialized algorithms are not widely used in the digital humanities and political science fields .
Approach: They propose a general neural framework based on topic models to enable customization of metadata.
Outcome: The proposed framework achieves strong performance with a manageable tradeoff between perplexity, coherence, and sparsity.
Are Neural Topic Models Broken? (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation paradigms are often divorced from real-world use . recent results have challenged the validity of the prevailing model evaluation paradigm .
Approach: They show that neural topic models fare worse in both respects compared to an established classical method.
Outcome: The proposed method outperforms the members of the ensemble in both respects.
Improving Entity Linking by Modeling Latent Relations between Mentions (P18-1)

Copied to clipboard

Challenge: Entity linking systems often exploit relations between textual mentions to decide if the linking decisions are compatible.
Approach: They treat relations as latent variables while optimizing the neural entity-linking model without supervision.
Outcome: The proposed model outperforms its relation-agnostic version and significantly outperformed its relational version.
Inflating Topic Relevance with Ideology: A Case Study of Political Ideology Bias in Social Topic Detection Models (2020.coling-main)

Copied to clipboard

Challenge: a study examines the impact of political ideology biases in training data . topic detection methods may contain or propagate certain biase resulting in a skewed data collection .
Approach: They propose to learn a text representation that is invariant to political ideology while still judging topic relevance.
Outcome: The proposed model can be invariant to political ideology while still judging topic relevance.
Stubborn Lexical Bias in Data and Models (2023.findings-acl)

Copied to clipboard

Challenge: Recent work has focused on spurious correlations between features and labels in training data . but, we find strong evidence of corresponding bias in the trained models .
Approach: They propose a method to reduce spurious correlations in training data by reweighting it using a large pool of extracted features.
Outcome: The proposed method reduces spurious correlations in training data, but still finds strong evidence of bias in trained models.
Boosting Entity Linking Performance by Leveraging Unlabeled Documents (P19-1)

Copied to clipboard

Challenge: a new approach to entity linking relies on unlabeled documents and Wikipedia . a supervised approach uses only natural information, such as unlabed documents .
Approach: They propose a method which exploits only naturally occurring information . they construct a high recall list of candidate entities for each mention in an unlabeled document .
Outcome: The proposed model outperforms fully-supervised state-of-the-art systems on standard test sets.
We Need to Measure Data Diversity in NLP — Better and Broader (2025.emnlp-main)

Copied to clipboard

Challenge: Language models exhibit remarkable natural language understanding and generation capabilities, but they have serious flaws, such as societal biases and spurious correlations.
Approach: They argue that interdisciplinary perspectives are essential for developing more fine-grained and valid measures of data diversity.
Outcome: The proposed measures are based on interdisciplinary perspectives and include a variety of datasets.
PhraseCTM: Correlated Topic Modeling on Phrases within Markov Random Fields (P18-2)

Copied to clipboard

Challenge: Recent phrase-level topic models are unable to capture the correlation structure among the discovered topics.
Approach: They propose a phrase-level topic model PhraseCTM and a method to find out the correlations of topics at phrase level.
Outcome: The proposed method shows that correlated topic modeling is a good way to interpret themes of corpus.
Tethering Broken Themes: Aligning Neural Topic Models with Labels and Authors (2025.findings-naacl)

Copied to clipboard

Challenge: Recent studies suggest that topic models do not align well with human intentions.
Approach: They propose a method to align neural topic models with both labels and authorship information.
Outcome: The proposed method improves existing models in terms of topic quality and alignment.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations