| Challenge: | a number of incidents of aggression and related events have increased over the web . the reach and extent of the Internet has given these events unprecedented power and influence to affect the lives of billions of people. |
| Approach: | They propose to develop an aggression tagset and an annotated corpus of Hindi-English code-mixed data from two of the most popular social networking / social media platforms in India -Twitter and Facebook. |
| Outcome: | The proposed dataset contains approximately 18k tweets and 21k facebook comments and is being released for further research in the field. |
Similar Papers
The ComMA Dataset V0.2: Annotating Aggression and Bias in Multilingual Social Media Discourse (2022.lrec-1)
Copied to clipboard
Ritesh Kumar, Shyam Ratan, Siddharth Singh, Enakshi Nandi, Laishram Niranjana Devi, Akash Bhagat, Yogesh Dawer, Bornini Lahiri, Akanksha Bansal, Atul Kr. Ojha
| Challenge: | 59,152 comments are annotated with a hierarchical, fine-grained taget marking aggression and bias of various kinds on social media platforms. |
| Approach: | They propose to annotate a multilingual dataset with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur. |
| Outcome: | The proposed dataset contains 59,152 comments in four languages, mostly code-mixed with English. |
Corpus Creation and Emotion Prediction for Hindi-English Code-Mixed Social Media Text (N18-4)
Copied to clipboard
| Challenge: | Emotion Prediction is a natural language processing task dealing with detection and classification of emotions in monolingual and bilingual texts. |
| Approach: | They propose a machine learning system which uses various machine learning techniques to detect emotion associated with tweets. |
| Outcome: | The proposed system uses various machine learning techniques to detect emotion associated with the text. |
Humor Detection in English-Hindi Code-Mixed Social Media Content : Corpus and Baseline System (L18-1)
Copied to clipboard
| Challenge: | a growing number of social media users are using code-mixing to detect humor . linguistics researchers are looking for methods to detect humorous content in text . |
| Approach: | They analyze a corpus of English-Hindi code-mixed tweets annotated with humorous(H) tags. |
| Outcome: | The proposed method detects humor in code-mixed tweets in English-Hindi. |
A Corpus of Turkish Offensive Language on Social Media (2020.lrec-1)
Copied to clipboard
| Challenge: | Identifying abusive, offensive, aggressive or in general inappropriate language has recently attracted interest of researchers from academic as well as commercial institutions. |
| Approach: | They propose to classify Turkish offensive language corpus using state-of-the-art annotation methods . they find 19 % of tweets contain some type of offensive language . |
| Outcome: | The proposed corpus of Turkish offensive language is the first of its kind in the world . the results show that 19 % of the tweets contain some type of offensive language . |
HindiMD: A Multi-domain Corpora for Low-resource Sentiment Analysis (2022.lrec-1)
Copied to clipboard
| Challenge: | Social media platforms such as Twitter and Facebook are a new channel of information dissemination for many negative groups for recruitment. |
| Approach: | They propose to use a social media sentiment analysis corpus annotated with the sentiment classes positive, negative and neutral to investigate the polarity of user-expressed opinions. |
| Outcome: | The proposed model is based on a set of benchmark datasets for sentiment analysis across a range of domains and languages. |
Improving the Detection of Multilingual Online Attacks with Rich Social Media Data from Singapore (2023.acl-long)
Copied to clipboard
Janosch Haber, Bertie Vidgen, Matthew Chapman, Vibhor Agarwal, Roy Ka-Wei Lee, Yong Keong Yap, Paul Röttger
| Challenge: | Toxic content is a global problem, but most resources for detecting toxic content are in English . new datasets and models for non-English languages focus exclusively on one language or dialect . |
| Approach: | They propose to use a multilingual dataset of online attacks to identify code-mixed toxic content in Singapore . they collect reddit comments in Indonesian, Malay, Singlish, and other languages and provide fine-grained hierarchical labels for attacks . |
| Outcome: | The proposed dataset provides fine-grained hierarchical labels for online attacks in Singapore . it shows that the metadata can be used for granular error analysis . |
Creation of Corpus and analysis in Code-Mixed Kannada-English Twitter data for Emotion Prediction (2020.coling-main)
Copied to clipboard
| Challenge: | Existing work on emotion prediction for resource-rich languages has focused on code-mixed social media corpus but not on Kannada-English code-mixed Twitter data. |
| Approach: | They analyze Kannada-English code-mixed Twitter corpus annotated with their respective ‘Emotion’ for each tweet. |
| Outcome: | The proposed model based on Kannada-English code-mixed Twitter corpus yielded an accuracy of 30% and 32% respectively. |
An Italian Twitter Corpus of Hate Speech against Immigrants (L18-1)
Copied to clipboard
| Challenge: | a recent study has annotated 6,000 tweets for hate speech against immigrants . the annotation scheme was designed to account for the multiplicity of factors that can contribute to the definition of a hate speech notion . |
| Approach: | They describe a Twitter corpus annotated for hate speech against immigrants . they propose a scheme that includes aggressiveness, offensiveness, irony, stereotype and intensity . |
| Outcome: | The proposed annotation scheme includes aggressiveness, offensiveness, irony, stereotype, intensity and (on an experimental basis) intensity. |
An Annotated Social Media Corpus for German (2020.lrec-1)
Copied to clipboard
| Challenge: | Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse. |
| Approach: | They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research. |
| Outcome: | The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets. |
A Dataset of Offensive Language in Kosovo Social Media (2022.lrec-1)
Copied to clipboard
| Challenge: | Social media are a central part of people’s lives but are rife with bullying and offensive language, creating an unsafe environment for their users. |
| Approach: | They propose to use user-generated comments on Facebook and YouTube from selected Kosovo news platforms to annotate offensive language in Albanian. |
| Outcome: | The proposed system improves on Danish but not Albanian, on offensive language recognition and distinguishing targeted and untargeted offence. |