| Challenge: | Vulgar words are employed in language use for several different functions, including expressing aggression, signaling group identity or the informality of the communication. |
| Approach: | They present a dataset of 7,800 tweets with six categories of vulgarity in which all instances of vulgar words are annotated with one of the six categories. |
| Outcome: | The proposed model can predict the category of a vulgar word based on the immediate context it appears in with 67.4 macro F1 across six classes. |
Similar Papers
Expressively vulgar: The socio-dynamics of vulgarity and its effects on sentiment analysis in social media (C18-1)
Copied to clipboard
| Challenge: | Vulgarity is a common linguistic expression and is used to perform several linguistic functions. |
| Approach: | They analyze vulgarity using tweets from users with known demographics and sentiment ratings for vulgar tweets to study sentiment analysis performance. |
| Outcome: | The proposed model can boost sentiment analysis performance by analyzing vulgar tweets and tweet sentiment ratings. |
Do You Really Want to Hurt Me? Predicting Abusive Swearing in Social Media (2020.lrec-1)
Copied to clipboard
| Challenge: | Swearing is a common form of verbal communication and occurs in social media and online forums . a study by a team of researchers has investigated the phenomenon of swearing in Twitter . |
| Approach: | They analyze tweets to determine abusive swearing using models that automatically predict it . they also investigate lexical, syntactic, and affective features that are more informative . |
| Outcome: | The proposed model can predict abusive swearing in a tweet context and provide an intrinsic evaluation of the model. |
Simple Models for Word Formation in Slang (N18-1)
Copied to clipboard
| Challenge: | slang is a popular vocabulary among young people due to its extragrammatical properties and the rise of social media. |
| Approach: | They propose a data-driven approach coupled with linguistic knowledge to develop generative models for three types of extra-grammatical word formation phenomena abounding in slang: Blends, Clippings, and Reduplicatives. |
| Outcome: | The proposed models show that slang exhibits extragrammatical properties that distinguish it from the standard form. |
How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Slang is a commonly used type of informal language that poses a daunting challenge to NLP systems. |
| Approach: | They compare human-attested slang and swiss-generated slurs with machine-generated ones . they find that LLMs have significant knowledge about the creative aspects of sling . |
| Outcome: | The proposed model compares human and machine-generated slang usages to find biases in human perceptions of sling . the results suggest that human-attested slms have significant knowledge about the creative aspects of a language . |
A Computational Framework for Slang Generation (2021.tacl-1)
Copied to clipboard
| Challenge: | Existing language models trained on large text corpora are biased toward formal language and under-represent slang. |
| Approach: | They propose a framework that models the speaker’s word choice in slang context by relating the conventional and sexist senses of a word while incorporating syntactic and contextual knowledge. |
| Outcome: | The proposed framework outperforms state-of-the-art language models and better predicts the historical emergence of slang word usages from 1960s to 2000s. |
A Computational Exploration of Pejorative Language in Social Media (2021.findings-emnlp)
Copied to clipboard
| Challenge: | In this paper, we examine the problem of pejorative language, an under-explored topic in computational linguistics. |
| Approach: | They propose to automatically disambiguate pejorative usage in social media . they leverage online dictionaries to build a multilingual lexicon of pejorativ terms . |
| Outcome: | The proposed model can automatically disambiguate pejorative usage in social media posts . the proposed model is based on dictionaries and tweets . |
Toward Informal Language Processing: Knowledge of Slang in Large Language Models (2024.naacl-long)
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have offered a strong potential for natural language systems to process informal language. |
| Approach: | They propose to use movie subtitles to evaluate slang in large language models . they find that smaller LLMs finetuned on the dataset achieve comparable performance . |
| Outcome: | The proposed dataset can be used to evaluate LLMs on slang detection and identification of regional and historical sources for interpretive insights. |
Toxic, Hateful, Offensive or Abusive? What Are We Really Classifying? An Empirical Analysis of Hate Speech Datasets (2020.lrec-1)
Copied to clipboard
| Challenge: | a recent study shows that many definitions are being used for equivalent concepts, making most datasets incompatible. |
| Approach: | They analyze six publicly available datasets to determine their similarity and compatibility . they propose to use Fast Text word vectors to analyze similarity between different datasets . |
| Outcome: | The proposed model performs better on similar datasets and worse on more non-offensive samples. |
A Survey of Meaning Representations – From Theory to Practical Utility (2024.naacl-long)
Copied to clipboard
| Challenge: | Symbolic meaning representations of natural language text have been studied since at least the 1960s . with the availability of large annotated corpora, the field has recently seen several new developments . |
| Approach: | They propose a framework for expressing meaning in natural language text using annotated corpora and a set of tools for machine learning. |
| Outcome: | The frameworks are based on a set of theoretical and practical problems and their applications. |
Predicting the Type and Target of Offensive Posts in Social Media (N19-1)
Copied to clipboard
| Challenge: | Prior work focused on detecting specific types of offensive content, such as hate speech, cyberbullying, or cyber-aggression. |
| Approach: | They propose to use a dataset to identify offensive content in social media . they compare the performance of different machine learning models to OLID . |
| Outcome: | The proposed dataset contains tweets annotated for offensive content using a fine-grained three-layer annotation scheme. |