Challenge: Vulgar words are employed in language use for several different functions, including expressing aggression, signaling group identity or the informality of the communication.
Approach: They present a dataset of 7,800 tweets with six categories of vulgarity in which all instances of vulgar words are annotated with one of the six categories.
Outcome: The proposed model can predict the category of a vulgar word based on the immediate context it appears in with 67.4 macro F1 across six classes.

Similar Papers

Expressively vulgar: The socio-dynamics of vulgarity and its effects on sentiment analysis in social media (C18-1)

Copied to clipboard

Challenge: Vulgarity is a common linguistic expression and is used to perform several linguistic functions.
Approach: They analyze vulgarity using tweets from users with known demographics and sentiment ratings for vulgar tweets to study sentiment analysis performance.
Outcome: The proposed model can boost sentiment analysis performance by analyzing vulgar tweets and tweet sentiment ratings.
Do You Really Want to Hurt Me? Predicting Abusive Swearing in Social Media (2020.lrec-1)

Copied to clipboard

Challenge: Swearing is a common form of verbal communication and occurs in social media and online forums . a study by a team of researchers has investigated the phenomenon of swearing in Twitter .
Approach: They analyze tweets to determine abusive swearing using models that automatically predict it . they also investigate lexical, syntactic, and affective features that are more informative .
Outcome: The proposed model can predict abusive swearing in a tweet context and provide an intrinsic evaluation of the model.
Simple Models for Word Formation in Slang (N18-1)

Copied to clipboard

Challenge: slang is a popular vocabulary among young people due to its extragrammatical properties and the rise of social media.
Approach: They propose a data-driven approach coupled with linguistic knowledge to develop generative models for three types of extra-grammatical word formation phenomena abounding in slang: Blends, Clippings, and Reduplicatives.
Outcome: The proposed models show that slang exhibits extragrammatical properties that distinguish it from the standard form.
How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages (2025.findings-emnlp)

Copied to clipboard

Challenge: Slang is a commonly used type of informal language that poses a daunting challenge to NLP systems.
Approach: They compare human-attested slang and swiss-generated slurs with machine-generated ones . they find that LLMs have significant knowledge about the creative aspects of sling .
Outcome: The proposed model compares human and machine-generated slang usages to find biases in human perceptions of sling . the results suggest that human-attested slms have significant knowledge about the creative aspects of a language .
A Computational Framework for Slang Generation (2021.tacl-1)

Copied to clipboard

Challenge: Existing language models trained on large text corpora are biased toward formal language and under-represent slang.
Approach: They propose a framework that models the speaker’s word choice in slang context by relating the conventional and sexist senses of a word while incorporating syntactic and contextual knowledge.
Outcome: The proposed framework outperforms state-of-the-art language models and better predicts the historical emergence of slang word usages from 1960s to 2000s.
A Computational Exploration of Pejorative Language in Social Media (2021.findings-emnlp)

Copied to clipboard

Challenge: In this paper, we examine the problem of pejorative language, an under-explored topic in computational linguistics.
Approach: They propose to automatically disambiguate pejorative usage in social media . they leverage online dictionaries to build a multilingual lexicon of pejorativ terms .
Outcome: The proposed model can automatically disambiguate pejorative usage in social media posts . the proposed model is based on dictionaries and tweets .
Toward Informal Language Processing: Knowledge of Slang in Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have offered a strong potential for natural language systems to process informal language.
Approach: They propose to use movie subtitles to evaluate slang in large language models . they find that smaller LLMs finetuned on the dataset achieve comparable performance .
Outcome: The proposed dataset can be used to evaluate LLMs on slang detection and identification of regional and historical sources for interpretive insights.
Toxic, Hateful, Offensive or Abusive? What Are We Really Classifying? An Empirical Analysis of Hate Speech Datasets (2020.lrec-1)

Copied to clipboard

Challenge: a recent study shows that many definitions are being used for equivalent concepts, making most datasets incompatible.
Approach: They analyze six publicly available datasets to determine their similarity and compatibility . they propose to use Fast Text word vectors to analyze similarity between different datasets .
Outcome: The proposed model performs better on similar datasets and worse on more non-offensive samples.
A Survey of Meaning Representations – From Theory to Practical Utility (2024.naacl-long)

Copied to clipboard

Challenge: Symbolic meaning representations of natural language text have been studied since at least the 1960s . with the availability of large annotated corpora, the field has recently seen several new developments .
Approach: They propose a framework for expressing meaning in natural language text using annotated corpora and a set of tools for machine learning.
Outcome: The frameworks are based on a set of theoretical and practical problems and their applications.
Predicting the Type and Target of Offensive Posts in Social Media (N19-1)

Copied to clipboard

Challenge: Prior work focused on detecting specific types of offensive content, such as hate speech, cyberbullying, or cyber-aggression.
Approach: They propose to use a dataset to identify offensive content in social media . they compare the performance of different machine learning models to OLID .
Outcome: The proposed dataset contains tweets annotated for offensive content using a fine-grained three-layer annotation scheme.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations