Leveraging Conflicts in Social Media Posts: Unintended Offense Dataset (2024.emnlp-main)
Copied to clipboard
| Challenge: | a new study examines the impact of conflict on multi-person communication datasets on offensive language . conflict datasets often neglect contextual information and focus on intended offenses . authors propose a conflict-based data collection method to analyze inter-conflict cues in multi-user communications . |
| Approach: | They propose a conflict-based data collection method to utilize inter-conflict cues in multi-person communications. |
| Outcome: | The proposed method improves the accuracy of detecting offensive language and enriches our understanding of conflict dynamics in digital communication. |
Similar Papers
Predicting the Type and Target of Offensive Posts in Social Media (N19-1)
Copied to clipboard
| Challenge: | Prior work focused on detecting specific types of offensive content, such as hate speech, cyberbullying, or cyber-aggression. |
| Approach: | They propose to use a dataset to identify offensive content in social media . they compare the performance of different machine learning models to OLID . |
| Outcome: | The proposed dataset contains tweets annotated for offensive content using a fine-grained three-layer annotation scheme. |
Conflicts in Texts: Data, Implications and Challenges (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Conflicts in data could reflect complexity of situations, changes that need to be explained and dealt with, difficulties in data annotation, and mistakes in generated outputs. |
| Approach: | This survey categorizes conflicting information into three key areas . they identify the areas where conflicting data can be ignored and undermine models' reliability and trustworthiness. |
| Outcome: | The findings highlight key challenges and future directions for developing conflict-aware NLP systems that can reason over and reconcile conflicting information more effectively. |
The ComMA Dataset V0.2: Annotating Aggression and Bias in Multilingual Social Media Discourse (2022.lrec-1)
Copied to clipboard
Ritesh Kumar, Shyam Ratan, Siddharth Singh, Enakshi Nandi, Laishram Niranjana Devi, Akash Bhagat, Yogesh Dawer, Bornini Lahiri, Akanksha Bansal, Atul Kr. Ojha
| Challenge: | 59,152 comments are annotated with a hierarchical, fine-grained taget marking aggression and bias of various kinds on social media platforms. |
| Approach: | They propose to annotate a multilingual dataset with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur. |
| Outcome: | The proposed dataset contains 59,152 comments in four languages, mostly code-mixed with English. |
Detecting and Reducing Bias in a High Stakes Domain (D19-1)
Copied to clipboard
| Challenge: | Existing research shows that a deep learning model can predict aggression and loss in posts by focusing on stop words such as “a” or “on”. |
| Approach: | They developed an approach to interpret a deep learning model that often bases its predictions on stop words such as "a" or "on" to tackle bias, they annotated the rationales and built models that drastically reduce bias. |
| Outcome: | The proposed model can predict aggression and loss in posts by using stop words such as "a" or "on" the new annotations enable us to quantitatively measure how justified the model predictions are, and build models that drastically reduce bias. |
Detecting Gang-Involved Escalation on Social Media Using Context (D18-1)
Copied to clipboard
Serina Chang, Ruiqi Zhong, Ethan Adams, Fei-Tzin Lee, Siddharth Varia, Desmond Patton, William Frey, Chris Kedzie, Kathy McKeown
| Challenge: | In cities such as Chicago, gang-involved youth have increasingly turned to social media to post about their experiences and intents online. |
| Approach: | They propose a system that uses domain-specific resources and contextual representations of the emotional and semantic content of the user’s recent tweets and their interactions with other users to detect Aggression and Loss in social media posts. |
| Outcome: | The proposed system improves on a large unlabeled dataset and incorporates contextual representations of the emotional and semantic content of the user’s recent tweets as well as their interactions with other users. |
Investigating Controversy Framing across Topics on Social Media (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a novel method for discovering framings of controversial problems is proposed . framers of controversial issues can be explored across topics, the paper argues . |
| Approach: | This paper proposes a method for discovering and articulating framing of controversial problems . framers offer valuable insights into how and why controversial problems are discussed online . |
| Outcome: | The proposed method enables the investigation of how controversy is framed across topics. |
Aggression-annotated Corpus of Hindi-English Code-mixed Data (L18-1)
Copied to clipboard
| Challenge: | a number of incidents of aggression and related events have increased over the web . the reach and extent of the Internet has given these events unprecedented power and influence to affect the lives of billions of people. |
| Approach: | They propose to develop an aggression tagset and an annotated corpus of Hindi-English code-mixed data from two of the most popular social networking / social media platforms in India -Twitter and Facebook. |
| Outcome: | The proposed dataset contains approximately 18k tweets and 21k facebook comments and is being released for further research in the field. |
CyberAgressionAdo-v1: a Dataset of Annotated Online Aggressions in French Collected through a Role-playing Game (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent studies have highlighted that private instant messaging platforms are major mediums of cyber aggression among teens. |
| Approach: | They present a dataset of aggressive chats in French collected through a role-playing game in high-schools . they provide insights on the different types of aggression and verbal abuse depending on the targeted victims . |
| Outcome: | The proposed dataset analyzes aggressive conversations in French on a role-playing game in high schools . it provides insights on the different types of aggression and verbal abuse depending on the targeted victims . |
Fighting Offensive Language on Social Media with Unsupervised Text Style Transfer (P18-2)
Copied to clipboard
| Challenge: | Existing methods to tackle the problem of offensive language in social media are based on machine learning. |
| Approach: | They propose a method for training encoder-decoders using non-parallel data . they use a collaborative classifier, attention and the cycle consistency loss . |
| Outcome: | The proposed method outperforms state-of-the-art text style transfer systems on Twitter and Reddit . it produces reliable non-offensive transferred sentences, the authors show . |
Dimensions of Online Conflict: Towards Modeling Agonism (2023.findings-emnlp)
Copied to clipboard
Matt Canute, Mali Jin, Hannah Holtzclaw, Alberto Lusoli, Philippa Adams, Mugdha Pandya, Maite Taboada, Diana Maynard, Wendy Hui Kyong Chun
| Challenge: | agonism fosters robust discussions, but hateful antagonism undermines constructive dialogue . a new study analyzes Twitter conversations to identify different dimensions of conflict . |
| Approach: | They annotated Twitter conversations related to trending controversial topics to model conflict on a richly annotized dataset. |
| Outcome: | The proposed model can help to moderate online conflicts and improve content monetization. |