Finding Microaggressions in the Wild: A Case for Locating Elusive Phenomena in Social Media Posts (D19-1)
Copied to clipboard
| Challenge: | Existing tools for hate speech detection and sentiment analysis cannot detect veiled offensiveness of microaggressions . linguistic subtlety of micro-aggressives has made it difficult to analyze their exact nature . |
| Approach: | They propose a typology of microaggressions based on a subset of data . they propose an objective criterion for annotation and an active-learning procedure . |
| Outcome: | The proposed typology of microaggressions is based on a subset of social media data. |
Similar Papers
Unifying Data Perspectivism and Personalization: An Application to Social Norms (2022.emnlp-main)
Copied to clipboard
| Challenge: | Obtaining a single ground truth is not possible or necessary for subjective tasks. |
| Approach: | They propose a set of personalization methods to model annotators and compare their effectiveness for predicting social norms. |
| Outcome: | The proposed model outperforms existing models and compares performance across subsets of social situations that vary by the closeness of the relationship between parties in conflict. |
Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media (2025.acl-long)
Copied to clipboard
| Challenge: | Social media platforms are experiencing a growing presence of AI-Generated Texts (AIGTs) however, the misuse of AIGTs could have profound implications for public opinion . |
| Approach: | They collect a dataset with 2.4M posts from 3 major social media platforms . they then construct a diverse dataset to train and evaluate AIGT detectors . |
| Outcome: | The proposed dataset analyzes 2.4M posts from 3 major social media platforms from 2022 to 2024 . it finds that Medium and Quora show marked increases in AAR . |
Classification without (Proper) Representation: Political Heterogeneity in Social Media and Its Implications for Classification and Behavioral Analysis (2022.findings-acl)
Copied to clipboard
| Challenge: | Prior work has shown that partisan leanings can be inferred from a diverse set of behavioral characteristics such as text, social networks, and even community participation. |
| Approach: | They test this assumption and show that commonly-used models do not generalize . they also show that political users are more toxic on the platform and inter-party interactions are even more toxic . |
| Outcome: | The proposed models do not generalize, indicating heterogeneous political users. |
#YouToo? Detection of Personal Recollections of Sexual Harassment on Social Media (P19-1)
Copied to clipboard
| Challenge: | a recent study has found that the disclosure of sexual abuse has positive psychological im- pacts. |
| Approach: | They propose to aggregate personal experiences of sexual harassment from Twitter posts to facilitate a better understanding of social media constructs and bring about social change. |
| Outcome: | The proposed model is compared with state-of-the-art models and is based on a three part Twitter-Specific Social Media Language Model. |
The ComMA Dataset V0.2: Annotating Aggression and Bias in Multilingual Social Media Discourse (2022.lrec-1)
Copied to clipboard
Ritesh Kumar, Shyam Ratan, Siddharth Singh, Enakshi Nandi, Laishram Niranjana Devi, Akash Bhagat, Yogesh Dawer, Bornini Lahiri, Akanksha Bansal, Atul Kr. Ojha
| Challenge: | 59,152 comments are annotated with a hierarchical, fine-grained taget marking aggression and bias of various kinds on social media platforms. |
| Approach: | They propose to annotate a multilingual dataset with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur. |
| Outcome: | The proposed dataset contains 59,152 comments in four languages, mostly code-mixed with English. |
SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference (D18-1)
Copied to clipboard
| Challenge: | a new dataset presents a task of grounded commonsense inference, unifying natural language inference and commonsensical reasoning. |
| Approach: | They propose a procedure that constructs a de-biased dataset by iteratively training stylistic classifiers and using them to filter the data. |
| Outcome: | The proposed procedure oversamples a de-biased dataset using state-of-the-art language models . human models struggle on the proposed procedure, indicating significant opportunities for future research. |
Breaking Down the Invisible Wall of Informal Fallacies in Online Discussions (2021.acl-long)
Copied to clipboard
| Challenge: | a number of people engage in unsound argumentation techniques to prove a claim on online platforms . fallacies are weak arguments that seem convincing, but their evidence does not prove or disprove the conclusion . |
| Approach: | They propose to use user comments containing fallacy mentions as noisy labels to classify fallacies . they use the pragma-dialectical theory of argumentation to study the most common fallacias on Reddit . |
| Outcome: | The proposed dataset of fallacies on reddit shows that neural models perform better in conversational context. |
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)
Copied to clipboard
Bolei Ma, Yuting Li, Wei Zhou, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, Barbara Plank
| Challenge: | linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions. |
| Approach: | They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications. |
| Outcome: | The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models . |
Author Profiling for Abuse Detection (C18-1)
Copied to clipboard
| Challenge: | Existing methods for detecting abusive content rely on textual cues and lexical cue information. |
| Approach: | They propose a method that incorporates community-based profiling features of Twitter users to detect abusive content by using a dataset of 16k tweets. |
| Outcome: | The proposed approach outperforms the current state-of-the-art in abuse detection on a dataset of 16k tweets. |
Speak up, Fight Back! Detection of Social Media Disclosures of Sexual Harassment (N19-3)
Copied to clipboard
| Challenge: | #MeToo movement provides platform to narrate personal experiences of sexual harassment. |
| Approach: | They propose a three-part ULMFiT architecture to tackle text subtleties in a classification task . they propose to annotate a manually annotated real-world dataset to test their approach . |
| Outcome: | The proposed model outperforms existing models that rely on handcrafted stylistic features and is more accurate than generic models. |