| Challenge: | Hate speech groups (HSGs) may negatively influence online platforms through their distinctive language, which may affect the tone and topic of discussion in other spaces if spread beyond the HSGs. |
| Approach: | They explore the linguistic style of the Manosphere on reddit and how it reflects their linguistic styles across communities. |
| Outcome: | The linguistic style of the Manosphere on Reddit is studied to determine whether it is harmful to health and community health. |
Similar Papers
Among Us: Language of Conspiracy Theorists on Mainstream Reddit (2026.acl-long)
Copied to clipboard
| Challenge: | Conspiracy theories are influential, alternative narratives that explain events through the actions of secretive, malevolent groups. |
| Approach: | They analyze a large-scale longitudinal dataset of over 500 million comments on reddit . they show that users exhibit distinctive linguistic patterns that enable machine learning models to distinguish them from the general population within individual communities. |
| Outcome: | The proposed model outperforms global classifiers by 17 percentage points. |
Stylistic approaches to predicting Reddit popularity in diglossia (2021.acl-srw)
Copied to clipboard
| Challenge: | Past research has shown that style is a strong predictor of community response, but what about a diglossia? |
| Approach: | They propose to use punctuation, stopwords and part-of-speech tags to predict the popularity of a Reddit post in a diglossia in Singapore where the basilect co-exists with an acrolect . |
| Outcome: | The proposed approach combines natural language processing (NLP) techniques with punctuation, stopwords and part-of-speech tags to predict popular posts in a diglossia. |
Ruddit: Norms of Offensiveness for English Reddit Comments (2021.acl-long)
Copied to clipboard
| Challenge: | Existing methods to detect offensive language have been limited by categorical labels . however, there are several challenges in the detection of such content . |
| Approach: | They analyze Reddit comments with fine-grained, real-valued offensiveness scores . they evaluate the ability of widely-used neural models to predict offensiveness . |
| Outcome: | The proposed method produces highly reliable offensiveness scores and can predict scores on reddit comments. |
Placing M-Phasis on the Plurality of Hate: A Feature-Based Corpus of Hate Online (2022.lrec-1)
Copied to clipboard
Dana Ruiter, Liane Reiners, Ashwin Geet D’Sa, Thomas Kleinbauer, Dominique Fohr, Irina Illina, Dietrich Klakow, Christian Schemer, Angeliki Monnier
| Challenge: | HS-related corpora over-simplify the phenomenon of hate by labelling user content with binary classes, e.g., hate/neutral . this ignores the complex and subjective nature of HS, which limits the real-life applicability of classifiers trained on these corporales. |
| Approach: | They present a corpus of 9k German and french user comments from migration-related news articles. |
| Outcome: | The proposed corpus is annotated with 23 features that become descriptors of various types of speech, ranging from critical comments to implicit and explicit expressions of hate. |
Whose Emotions and Moral Sentiments do Language Models Reflect? (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing research has focused on positional alignment, which measures how closely the models mimic the opinions and stances of different social groups. |
| Approach: | They define the problem of affective alignment, which measures how LMs’ emotional and moral tone represents those of different groups. |
| Outcome: | The results show that the models represent the perspectives of some social groups better than others, suggesting a systemic bias within LMs. |
Do language models practice what they preach? Examining language ideologies about gendered language reform encoded in LLMs (2025.coling-main)
Copied to clipboard
| Challenge: | Language ideologies are evaluative ideas or beliefs about language, such as ideas about what is "correct", "natural" or "articulate". |
| Approach: | They use gender-neutral variants more often when more explicit metalinguistic context is provided. |
| Outcome: | The findings show that language ideologies in LLMs can vary, which may be unexpected to users. |
Classification without (Proper) Representation: Political Heterogeneity in Social Media and Its Implications for Classification and Behavioral Analysis (2022.findings-acl)
Copied to clipboard
| Challenge: | Prior work has shown that partisan leanings can be inferred from a diverse set of behavioral characteristics such as text, social networks, and even community participation. |
| Approach: | They test this assumption and show that commonly-used models do not generalize . they also show that political users are more toxic on the platform and inter-party interactions are even more toxic . |
| Outcome: | The proposed models do not generalize, indicating heterogeneous political users. |
Do Large Language Models Understand Mansplaining? Well, Actually... (2024.lrec-main)
Copied to clipboard
| Challenge: | Gender bias has been studied by the NLP community, but other variations of it, such as mansplaining, have received little attention. |
| Approach: | They propose to analyze a corpus of 886 mansplaining stories experienced by women and examine how Large Language Models can understand and identify mansplaiting. |
| Outcome: | The proposed models reproduce some of the social patterns behind mansplaining situations by praising men for giving unsolicited advice to women. |
Exploring the Impact of Language Switching on Personality Traits in LLMs (2025.coling-main)
Copied to clipboard
| Challenge: | Using three personality tests, we examine the extent to which LLMs align with humans when personality shifts are associated with language changes. |
| Approach: | They propose to use the Eysenck Personality Questionnaire-Revised to examine whether LLMs align with humans when personality shifts are associated with language changes. |
| Outcome: | The results show that language-switching affects personality traits in multilingual individuals, and that it is not translation-related. |
Splits! Flexible Sociocultural Linguistic Investigation at Scale (2026.acl-long)
Copied to clipboard
| Challenge: | Variation in language use offers a rich lens into cultural perspectives, values, and opinions. |
| Approach: | They propose to construct a "sandbox" for systematic and flexible sociolinguistic research by splitting a reddit dataset into demographically/topically split SLPs. |
| Outcome: | The proposed method analyzes a demographically/topically split Reddit dataset validated by self-identification and replicating several known SLPs from existing literature. |