Challenge: Existing tools for hate speech detection and sentiment analysis cannot detect veiled offensiveness of microaggressions . linguistic subtlety of micro-aggressives has made it difficult to analyze their exact nature .
Approach: They propose a typology of microaggressions based on a subset of data . they propose an objective criterion for annotation and an active-learning procedure .
Outcome: The proposed typology of microaggressions is based on a subset of social media data.

Similar Papers

Unifying Data Perspectivism and Personalization: An Application to Social Norms (2022.emnlp-main)

Copied to clipboard

Challenge: Obtaining a single ground truth is not possible or necessary for subjective tasks.
Approach: They propose a set of personalization methods to model annotators and compare their effectiveness for predicting social norms.
Outcome: The proposed model outperforms existing models and compares performance across subsets of social situations that vary by the closeness of the relationship between parties in conflict.
Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media (2025.acl-long)

Copied to clipboard

Challenge: Social media platforms are experiencing a growing presence of AI-Generated Texts (AIGTs) however, the misuse of AIGTs could have profound implications for public opinion .
Approach: They collect a dataset with 2.4M posts from 3 major social media platforms . they then construct a diverse dataset to train and evaluate AIGT detectors .
Outcome: The proposed dataset analyzes 2.4M posts from 3 major social media platforms from 2022 to 2024 . it finds that Medium and Quora show marked increases in AAR .
Classification without (Proper) Representation: Political Heterogeneity in Social Media and Its Implications for Classification and Behavioral Analysis (2022.findings-acl)

Copied to clipboard

Challenge: Prior work has shown that partisan leanings can be inferred from a diverse set of behavioral characteristics such as text, social networks, and even community participation.
Approach: They test this assumption and show that commonly-used models do not generalize . they also show that political users are more toxic on the platform and inter-party interactions are even more toxic .
Outcome: The proposed models do not generalize, indicating heterogeneous political users.
#YouToo? Detection of Personal Recollections of Sexual Harassment on Social Media (P19-1)

Copied to clipboard

Challenge: a recent study has found that the disclosure of sexual abuse has positive psychological im- pacts.
Approach: They propose to aggregate personal experiences of sexual harassment from Twitter posts to facilitate a better understanding of social media constructs and bring about social change.
Outcome: The proposed model is compared with state-of-the-art models and is based on a three part Twitter-Specific Social Media Language Model.
The ComMA Dataset V0.2: Annotating Aggression and Bias in Multilingual Social Media Discourse (2022.lrec-1)

Copied to clipboard

Challenge: 59,152 comments are annotated with a hierarchical, fine-grained taget marking aggression and bias of various kinds on social media platforms.
Approach: They propose to annotate a multilingual dataset with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur.
Outcome: The proposed dataset contains 59,152 comments in four languages, mostly code-mixed with English.
SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference (D18-1)

Copied to clipboard

Challenge: a new dataset presents a task of grounded commonsense inference, unifying natural language inference and commonsensical reasoning.
Approach: They propose a procedure that constructs a de-biased dataset by iteratively training stylistic classifiers and using them to filter the data.
Outcome: The proposed procedure oversamples a de-biased dataset using state-of-the-art language models . human models struggle on the proposed procedure, indicating significant opportunities for future research.
Breaking Down the Invisible Wall of Informal Fallacies in Online Discussions (2021.acl-long)

Copied to clipboard

Challenge: a number of people engage in unsound argumentation techniques to prove a claim on online platforms . fallacies are weak arguments that seem convincing, but their evidence does not prove or disprove the conclusion .
Approach: They propose to use user comments containing fallacy mentions as noisy labels to classify fallacies . they use the pragma-dialectical theory of argumentation to study the most common fallacias on Reddit .
Outcome: The proposed dataset of fallacies on reddit shows that neural models perform better in conversational context.
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)

Copied to clipboard

Challenge: linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions.
Approach: They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications.
Outcome: The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models .
Author Profiling for Abuse Detection (C18-1)

Copied to clipboard

Challenge: Existing methods for detecting abusive content rely on textual cues and lexical cue information.
Approach: They propose a method that incorporates community-based profiling features of Twitter users to detect abusive content by using a dataset of 16k tweets.
Outcome: The proposed approach outperforms the current state-of-the-art in abuse detection on a dataset of 16k tweets.
Speak up, Fight Back! Detection of Social Media Disclosures of Sexual Harassment (N19-3)

Copied to clipboard

Challenge: #MeToo movement provides platform to narrate personal experiences of sexual harassment.
Approach: They propose a three-part ULMFiT architecture to tackle text subtleties in a classification task . they propose to annotate a manually annotated real-world dataset to test their approach .
Outcome: The proposed model outperforms existing models that rely on handcrafted stylistic features and is more accurate than generic models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations