Style-Shifting Behaviour of the Manosphere on Reddit (2024.emnlp-main)

Copied to clipboard

Challenge: Hate speech groups (HSGs) may negatively influence online platforms through their distinctive language, which may affect the tone and topic of discussion in other spaces if spread beyond the HSGs.
Approach: They explore the linguistic style of the Manosphere on reddit and how it reflects their linguistic styles across communities.
Outcome: The linguistic style of the Manosphere on Reddit is studied to determine whether it is harmful to health and community health.

Similar Papers

Among Us: Language of Conspiracy Theorists on Mainstream Reddit (2026.acl-long)

Copied to clipboard

Challenge: Conspiracy theories are influential, alternative narratives that explain events through the actions of secretive, malevolent groups.
Approach: They analyze a large-scale longitudinal dataset of over 500 million comments on reddit . they show that users exhibit distinctive linguistic patterns that enable machine learning models to distinguish them from the general population within individual communities.
Outcome: The proposed model outperforms global classifiers by 17 percentage points.
Stylistic approaches to predicting Reddit popularity in diglossia (2021.acl-srw)

Copied to clipboard

Challenge: Past research has shown that style is a strong predictor of community response, but what about a diglossia?
Approach: They propose to use punctuation, stopwords and part-of-speech tags to predict the popularity of a Reddit post in a diglossia in Singapore where the basilect co-exists with an acrolect .
Outcome: The proposed approach combines natural language processing (NLP) techniques with punctuation, stopwords and part-of-speech tags to predict popular posts in a diglossia.
Ruddit: Norms of Offensiveness for English Reddit Comments (2021.acl-long)

Copied to clipboard

Challenge: Existing methods to detect offensive language have been limited by categorical labels . however, there are several challenges in the detection of such content .
Approach: They analyze Reddit comments with fine-grained, real-valued offensiveness scores . they evaluate the ability of widely-used neural models to predict offensiveness .
Outcome: The proposed method produces highly reliable offensiveness scores and can predict scores on reddit comments.
Placing M-Phasis on the Plurality of Hate: A Feature-Based Corpus of Hate Online (2022.lrec-1)

Copied to clipboard

Challenge: HS-related corpora over-simplify the phenomenon of hate by labelling user content with binary classes, e.g., hate/neutral . this ignores the complex and subjective nature of HS, which limits the real-life applicability of classifiers trained on these corporales.
Approach: They present a corpus of 9k German and french user comments from migration-related news articles.
Outcome: The proposed corpus is annotated with 23 features that become descriptors of various types of speech, ranging from critical comments to implicit and explicit expressions of hate.
Whose Emotions and Moral Sentiments do Language Models Reflect? (2024.findings-acl)

Copied to clipboard

Challenge: Existing research has focused on positional alignment, which measures how closely the models mimic the opinions and stances of different social groups.
Approach: They define the problem of affective alignment, which measures how LMs’ emotional and moral tone represents those of different groups.
Outcome: The results show that the models represent the perspectives of some social groups better than others, suggesting a systemic bias within LMs.
Do language models practice what they preach? Examining language ideologies about gendered language reform encoded in LLMs (2025.coling-main)

Copied to clipboard

Challenge: Language ideologies are evaluative ideas or beliefs about language, such as ideas about what is "correct", "natural" or "articulate".
Approach: They use gender-neutral variants more often when more explicit metalinguistic context is provided.
Outcome: The findings show that language ideologies in LLMs can vary, which may be unexpected to users.
Classification without (Proper) Representation: Political Heterogeneity in Social Media and Its Implications for Classification and Behavioral Analysis (2022.findings-acl)

Copied to clipboard

Challenge: Prior work has shown that partisan leanings can be inferred from a diverse set of behavioral characteristics such as text, social networks, and even community participation.
Approach: They test this assumption and show that commonly-used models do not generalize . they also show that political users are more toxic on the platform and inter-party interactions are even more toxic .
Outcome: The proposed models do not generalize, indicating heterogeneous political users.
Do Large Language Models Understand Mansplaining? Well, Actually... (2024.lrec-main)

Copied to clipboard

Challenge: Gender bias has been studied by the NLP community, but other variations of it, such as mansplaining, have received little attention.
Approach: They propose to analyze a corpus of 886 mansplaining stories experienced by women and examine how Large Language Models can understand and identify mansplaiting.
Outcome: The proposed models reproduce some of the social patterns behind mansplaining situations by praising men for giving unsolicited advice to women.
Exploring the Impact of Language Switching on Personality Traits in LLMs (2025.coling-main)

Copied to clipboard

Challenge: Using three personality tests, we examine the extent to which LLMs align with humans when personality shifts are associated with language changes.
Approach: They propose to use the Eysenck Personality Questionnaire-Revised to examine whether LLMs align with humans when personality shifts are associated with language changes.
Outcome: The results show that language-switching affects personality traits in multilingual individuals, and that it is not translation-related.
Splits! Flexible Sociocultural Linguistic Investigation at Scale (2026.acl-long)

Copied to clipboard

Challenge: Variation in language use offers a rich lens into cultural perspectives, values, and opinions.
Approach: They propose to construct a "sandbox" for systematic and flexible sociolinguistic research by splitting a reddit dataset into demographically/topically split SLPs.
Outcome: The proposed method analyzes a demographically/topically split Reddit dataset validated by self-identification and replicating several known SLPs from existing literature.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations