Mapping (Dis-)Information Flow about the MH17 Plane Crash (D19-50)

Copied to clipboard

Challenge: Digital media enables fast sharing of information, but also disinformation . studies on the spread of disinformation on social media focused on small, manually annotated datasets or used proxys for data annotation.
Approach: They propose to use text classifiers to label Twitter content related to the MH17 crash to improve annotation accuracy.
Outcome: The proposed classifier improves over a hashtag-based baseline, but still remains a challenge in labelling pro-Russian and pro-Ukrainian content with high precision.

Similar Papers

Sentiment Analysis: It’s Complicated! (N18-1)

Copied to clipboard

Challenge: a dataset of over 7,000 tweets annotated with 5x coverage is used for sentiment analysis . a "complicated" class of sentiment is used to categorize text based on a predefined notion of sentiment .
Approach: They propose to use a "complicated" class of sentiment to categorize tweets . they build a publicly available tweet sentiment analysis dataset .
Outcome: The proposed classifiers perform better over a new publicly available TSA dataset . the classifier performance is compared with existing methods and improves on existing ones .
MiDe22: An Annotated Multi-Event Tweet Dataset for Misinformation Detection (2024.lrec-main)

Copied to clipboard

Challenge: a new dataset of misinformation labels is being developed to detect misinformation on social media platforms . misinformation is spread in many domains including but not limited to health, politics, and disasters .
Approach: They construct a dataset of 5,284 English and 5,064 Turkish tweets with misinformation labels . they use the dataset to analyze misinformation spread and to evaluate misinformation detection .
Outcome: The proposed dataset includes 5,284 English and 5,064 Turkish tweets with misinformation labels for several recent events between 2020 and 2022.
Event-Related Bias Removal for Real-time Disaster Events (2020.findings-emnlp)

Copied to clipboard

Challenge: Social media has become an important tool to share information about crisis events such as natural disasters and mass attacks.
Approach: They propose to train an adversarial neural model to remove latent event-specific biases and improve the performance on tweet importance classification.
Outcome: The proposed model removes event-specific biases and improves on tweet importance classification.
Unifying Data Perspectivism and Personalization: An Application to Social Norms (2022.emnlp-main)

Copied to clipboard

Challenge: Obtaining a single ground truth is not possible or necessary for subjective tasks.
Approach: They propose a set of personalization methods to model annotators and compare their effectiveness for predicting social norms.
Outcome: The proposed model outperforms existing models and compares performance across subsets of social situations that vary by the closeness of the relationship between parties in conflict.
An Expert Annotated Dataset for the Detection of Online Misogyny (2021.eacl-main)

Copied to clipboard

Challenge: Existing studies have found that misogynistic content is pervasive on some Reddit communities, but a training dataset for misogorical classification has not been created with the data.
Approach: They propose a hierarchical taxonomy and an expert labelled dataset to enable automatic classification of online misogynistic content.
Outcome: The proposed taxonomy and an expert labelled dataset are made freely available for future research.
The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection (2025.findings-naacl)

Copied to clipboard

Challenge: Recent research suggests using Large Language Models (LLMs) to automate the annotation process, reducing these costs while maintaining data quality.
Approach: They propose to use Large Language Models to automate annotation process and train classifiers on large datasets.
Outcome: The proposed model outperforms all of the annotator LLMs on two media bias benchmark datasets (BABE and BASIL) while maintaining data quality.
Categorizing and Inferring the Relationship between the Text and Image of Twitter Posts (P19-1)

Copied to clipboard

Challenge: Social media posts often contain images to provide content, provide context, or express feelings.
Approach: They build and release a dataset of image tweets annotated with four different classes which express whether the text or the image provides additional information to the other modality.
Outcome: The proposed method can be used in several downstream applications including pre-training image tagging models and collecting distantly supervised data for image captioning.
Understanding Emotions: A Dataset of Tweets to Study Interactions between Affect Categories (L18-1)

Copied to clipboard

Challenge: a new dataset is used to classify text into positive, negative, and neutral classes . a large amount of work on automatic detecting emotions from text has focused on classifying text into basic emotion categories .
Approach: They use Twitter as the source of the textual data they annotate to find out which emotions often present together in tweets .
Outcome: The proposed dataset is useful for training and testing supervised machine learning algorithms . it is based on the results of the SemEval-2018 task 1: Affect in Tweets .
TWEETSPIN: Fine-grained Propaganda Detection in Social Media Using Multi-View Representations (2022.naacl-main)

Copied to clipboard

Challenge: Recent studies on propaganda detection involve document and fragment-level analyses of news articles.
Approach: They propose a neural approach to detect and categorize propaganda tweets across fine-grained categories . they use a dataset containing tweets weakly annotated with different propaganda techniques .
Outcome: The proposed method outperforms benchmark methods and transfers knowledge to low-resource news domains.
RuSentiment: An Enriched Sentiment Analysis Dataset for Social Media in Russian (C18-1)

Copied to clipboard

Challenge: RuSentiment is currently the largest in its class for Russian, with 31,185 posts annotated with Fleiss’ kappa of 0.58 (3 annotations per post).
Approach: They propose to use RuSentiment to annotate social media posts in Russian with a kappa of 0.58 and a set of annotation guidelines that are extensible to other languages.
Outcome: The proposed dataset is the largest in its class for Russian, with 31,185 posts annotated with Fleiss’ kappa of 0.58 (3 annotations per post).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations