Challenge: Existing methods to classify social media posts into topics have been used to class up documents into topics.
Approach: They propose a neural model that automatically associates social media posts with topics to solve these challenges.
Outcome: The proposed model outperforms existing methods in the context of Twitter where the topic space is 10 times larger with potentially multiple topic associations per Tweet.

Similar Papers

Twitter Topic Classification (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to identify topics from posts are difficult to interpret and can differ from corpus to corpus.
Approach: They propose a task based on tweet topic classification and release two datasets that can be used to train and test models.
Outcome: The proposed task is based on two datasets from recent time periods and provides training and testing data.
Neural Multimodal Topic Modeling: A Comprehensive Evaluation (2024.lrec-main)

Copied to clipboard

Challenge: Neural topic models can find coherent and diverse topics in textual data, but they are limited in dealing with multimodal datasets.
Approach: They propose two new topic modeling solutions and two new evaluation metrics for document multimodality.
Outcome: The proposed models generate coherent and diverse topics on a rich dataset.
Multilingual Topic Classification in X: Dataset and Analysis (2024.emnlp-main)

Copied to clipboard

Challenge: Social media platforms such as X (Twitter), Snapchat and Instagram provide an environment for content creation and information sharing.
Approach: They propose a multilingual dataset featuring tweet topic classification in four languages . they leverage X-Topic to perform cross-linguistic and multilingual analysis .
Outcome: The proposed dataset includes topics in four languages and is useful for cross-linguistic analysis and the development of robust multilingual models.
Topic Modeling for Short Texts with Large Language Models (2024.acl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) can be used to solve topic modeling challenges for short texts by contextually learning the meanings of words.
Approach: They propose two approaches to using Large Language Models (LLMs) for topic modeling: parallel prompting and sequential prompting.
Outcome: The proposed methods identify more coherent topics than existing ones while maintaining the diversity of the induced topics.
Multi-source Neural Topic Modeling in Multi-view Embedding Spaces (2021.naacl-main)

Copied to clipboard

Challenge: Recent work has used pre-trained word embeddings to address data sparsity in short-text or small document collections.
Approach: They propose a neural topic modeling framework using multi-view embedding spaces to improve topic quality and deal with polysemy.
Outcome: The proposed framework improves topic quality and deal with polysemy.
Improving the TENOR of Labeling: Re-evaluating Topic Models for Content Analysis (2024.eacl-long)

Copied to clipboard

Challenge: Existing evaluation metrics such as coherence and coherency are inadequate for neural topic models.
Approach: They conduct the first evaluation of neural, supervised and classical topic models in an interactive task-based setting.
Outcome: The proposed model performs better on cluster evaluation metrics and human evaluations than classical models on real-world tasks.
TWEETQA: A Social Media Focused Question Answering Dataset (P19-1)

Copied to clipboard

Challenge: Social media is becoming an important realtime information source, especially during natural disasters and emergencies.
Approach: They present a large-scale dataset for question answering over social media data . they gather tweets used by journalists and ask human annotators to write questions upon them .
Outcome: The proposed dataset shows that neural models that perform well on formal texts are limited in their performance . the proposed model is still lagging behind human performance with a large margin .
Revisiting Automated Topic Model Evaluation with Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Topic models are an unsupervised dimensionality reduction technique that help organize large text collections.
Approach: They propose to use large language models to evaluate document output and determine optimal number of topics.
Outcome: The proposed model performs better on coherence ratings of word sets than on intrustion detection.
Self-Supervised Neural Topic Modeling (2021.findings-emnlp)

Copied to clipboard

Challenge: Topic models are useful tools for analyzing and interpreting the main underlying themes of large corpora of text.
Approach: They propose a self-supervised neural topic model that learns a topic representation jointly from three co-occurring words and a document that the triple originates from.
Outcome: The proposed model outperforms existing topic models in coherence metrics and document clustering accuracy.
Topics as Entity Clusters: Entity-based Topics from Large Language Models and Graph Neural Networks (2024.lrec-main)

Copied to clipboard

Challenge: Topic models aim to reveal latent structures within corpus of text through term-frequency statistics over bag-of-words representations.
Approach: They propose to use bimodal vector representations of entities to extract latent representations from large language models and graph neural networks trained on symbolic relations to derive the most salient aspects of these conceptual units.
Outcome: The proposed approach is better suited to working with entities than state-of-the-art models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations