Aggression-annotated Corpus of Hindi-English Code-mixed Data (L18-1)

Copied to clipboard

Challenge: a number of incidents of aggression and related events have increased over the web . the reach and extent of the Internet has given these events unprecedented power and influence to affect the lives of billions of people.
Approach: They propose to develop an aggression tagset and an annotated corpus of Hindi-English code-mixed data from two of the most popular social networking / social media platforms in India -Twitter and Facebook.
Outcome: The proposed dataset contains approximately 18k tweets and 21k facebook comments and is being released for further research in the field.

Similar Papers

The ComMA Dataset V0.2: Annotating Aggression and Bias in Multilingual Social Media Discourse (2022.lrec-1)

Copied to clipboard

Challenge: 59,152 comments are annotated with a hierarchical, fine-grained taget marking aggression and bias of various kinds on social media platforms.
Approach: They propose to annotate a multilingual dataset with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur.
Outcome: The proposed dataset contains 59,152 comments in four languages, mostly code-mixed with English.
Corpus Creation and Emotion Prediction for Hindi-English Code-Mixed Social Media Text (N18-4)

Copied to clipboard

Challenge: Emotion Prediction is a natural language processing task dealing with detection and classification of emotions in monolingual and bilingual texts.
Approach: They propose a machine learning system which uses various machine learning techniques to detect emotion associated with tweets.
Outcome: The proposed system uses various machine learning techniques to detect emotion associated with the text.
Humor Detection in English-Hindi Code-Mixed Social Media Content : Corpus and Baseline System (L18-1)

Copied to clipboard

Challenge: a growing number of social media users are using code-mixing to detect humor . linguistics researchers are looking for methods to detect humorous content in text .
Approach: They analyze a corpus of English-Hindi code-mixed tweets annotated with humorous(H) tags.
Outcome: The proposed method detects humor in code-mixed tweets in English-Hindi.
A Corpus of Turkish Offensive Language on Social Media (2020.lrec-1)

Copied to clipboard

Challenge: Identifying abusive, offensive, aggressive or in general inappropriate language has recently attracted interest of researchers from academic as well as commercial institutions.
Approach: They propose to classify Turkish offensive language corpus using state-of-the-art annotation methods . they find 19 % of tweets contain some type of offensive language .
Outcome: The proposed corpus of Turkish offensive language is the first of its kind in the world . the results show that 19 % of the tweets contain some type of offensive language .
HindiMD: A Multi-domain Corpora for Low-resource Sentiment Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Social media platforms such as Twitter and Facebook are a new channel of information dissemination for many negative groups for recruitment.
Approach: They propose to use a social media sentiment analysis corpus annotated with the sentiment classes positive, negative and neutral to investigate the polarity of user-expressed opinions.
Outcome: The proposed model is based on a set of benchmark datasets for sentiment analysis across a range of domains and languages.
Improving the Detection of Multilingual Online Attacks with Rich Social Media Data from Singapore (2023.acl-long)

Copied to clipboard

Challenge: Toxic content is a global problem, but most resources for detecting toxic content are in English . new datasets and models for non-English languages focus exclusively on one language or dialect .
Approach: They propose to use a multilingual dataset of online attacks to identify code-mixed toxic content in Singapore . they collect reddit comments in Indonesian, Malay, Singlish, and other languages and provide fine-grained hierarchical labels for attacks .
Outcome: The proposed dataset provides fine-grained hierarchical labels for online attacks in Singapore . it shows that the metadata can be used for granular error analysis .
Creation of Corpus and analysis in Code-Mixed Kannada-English Twitter data for Emotion Prediction (2020.coling-main)

Copied to clipboard

Challenge: Existing work on emotion prediction for resource-rich languages has focused on code-mixed social media corpus but not on Kannada-English code-mixed Twitter data.
Approach: They analyze Kannada-English code-mixed Twitter corpus annotated with their respective ‘Emotion’ for each tweet.
Outcome: The proposed model based on Kannada-English code-mixed Twitter corpus yielded an accuracy of 30% and 32% respectively.
An Italian Twitter Corpus of Hate Speech against Immigrants (L18-1)

Copied to clipboard

Challenge: a recent study has annotated 6,000 tweets for hate speech against immigrants . the annotation scheme was designed to account for the multiplicity of factors that can contribute to the definition of a hate speech notion .
Approach: They describe a Twitter corpus annotated for hate speech against immigrants . they propose a scheme that includes aggressiveness, offensiveness, irony, stereotype and intensity .
Outcome: The proposed annotation scheme includes aggressiveness, offensiveness, irony, stereotype, intensity and (on an experimental basis) intensity.
An Annotated Social Media Corpus for German (2020.lrec-1)

Copied to clipboard

Challenge: Hate Speech (HS) against ethnic, religious and national minorities is a growing concern in online discourse.
Approach: They present the German Twitter section of a large (2 billion word) bilingual Social Media corpus for Hate Speech research.
Outcome: The proposed parser achieved F-scores of 97% for morphology and 92% for syntax on a cross-section of tweets.
A Dataset of Offensive Language in Kosovo Social Media (2022.lrec-1)

Copied to clipboard

Challenge: Social media are a central part of people’s lives but are rife with bullying and offensive language, creating an unsafe environment for their users.
Approach: They propose to use user-generated comments on Facebook and YouTube from selected Kosovo news platforms to annotate offensive language in Albanian.
Outcome: The proposed system improves on Danish but not Albanian, on offensive language recognition and distinguishing targeted and untargeted offence.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations