Papers with Reddit

108 papers
Grounding in social media: An approach to building a chit-chat dialogue model (2022.naacl-srw)

Copied to clipboard

Challenge: Existing open-domain dialogue models fail to capture and utilize external knowledge, leading to repetitive or generic responses to unseen utterances.
Approach: They propose to use social media comments to improve the raw conversation ability of open-domain dialogue systems.
Outcome: The proposed model improves the raw conversation ability of open-domain dialogue systems by mimicking human responses through casual interactions found on social media.
Semantic shift in social networks (2021.starsem-1)

Copied to clipboard

Challenge: lexical semantic change manifests differently across different communities, according to a new study . social network analysis is a tool of sociolinguists studying variation and change .
Approach: They use distributional methods to quantify lexical semantic change and induce a social network on communities based on interactions between members.
Outcome: The proposed method is based on interactions between members and the community.
Tkol, Httt, and r/radiohead: High Affinity Terms in Reddit Communities (D19-55)

Copied to clipboard

Challenge: Language is an important marker of a cultural group, large or small.
Approach: They analyze the evolution of high affinity terms across 2600 subreddits . they show that high affinity words are effective signals of loyal communities .
Outcome: The results show that high affinity terms are effective signals of loyal communities, they undergo more semantic shift than low affinity terms, and they are partial barrier to entry for new users.
Attentive Interaction Model: Modeling Changes in View in Argumentation (N18-1)

Copied to clipboard

Challenge: Prior work on argumentation in the NLP community has focused mainly on the first goal and has missed more nuanced and complex details of viewpoints.
Approach: They propose a neural architecture that explicitly models the interplay between an Opinion Holder's (OH's) reasoning and a challenger's argument to predict if the argument succeeded in altering the OH' s view.
Outcome: The proposed model outperforms several baselines on discussions on the Change My View forum on Reddit.
Neural Argument Generation Augmented with Externally Retrieved Evidence (P18-1)

Copied to clipboard

Challenge: Existing methods for generating arguments are limited to retrieval-based methods.
Approach: They propose an encoder-decoder-based argument generation model enriched with externally retrieved evidence from Wikipedia.
Outcome: The proposed model generates arguments with more topic-relevant content than current models based on automatic evaluation and human assessments on a large-scale dataset from reddit.
‘Am I the Bad One’? Predicting the Moral Judgement of the Crowd Using Pre–trained Language Models (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies on NLP touch upon moral contexts in text.
Approach: They construct a dataset that can be used for moral judgement tasks on a popular reddit subreddit.
Outcome: The proposed model passes moral judgements on posts from a popular reddit subreddit . it shows that the model can be fine tuned and improves across the datasets .
Identifying and Understanding User Reactions to Deceptive and Trusted Social News Sources (P18-2)

Copied to clipboard

Challenge: a new study examines how users react to news sources with different levels of credibility . a recent study found that 59% of bitly-URLs on Twitter are shared without ever being read .
Approach: They develop a model to classify user reactions into one of nine types . they also measure the speed and type of reaction for trusted and deceptive news sources .
Outcome: The proposed model classifies user reactions into one of nine types, such as answer, elaboration, and question, etc.
Fighting Offensive Language on Social Media with Unsupervised Text Style Transfer (P18-2)

Copied to clipboard

Challenge: Existing methods to tackle the problem of offensive language in social media are based on machine learning.
Approach: They propose a method for training encoder-decoders using non-parallel data . they use a collaborative classifier, attention and the cycle consistency loss .
Outcome: The proposed method outperforms state-of-the-art text style transfer systems on Twitter and Reddit . it produces reliable non-offensive transferred sentences, the authors show .
Deep Dirichlet Multinomial Regression (N18-1)

Copied to clipboard

Challenge: supervised topic models can incorporate arbitrary document-level features to inform topic priors, but their ability to model corpora is limited by the representation and selection of these features.
Approach: They propose a generative topic model that simultaneously learns document feature representations and topics.
Outcome: The proposed model outperforms DMR and LDA on three datasets and human subjects judge it more representative of associated document features.
MetaMeme: A Dataset for Meme Template and Meta-Category Classification (2025.naacl-srw)

Copied to clipboard

Challenge: a new dataset for classifying memes by their template and communicative intent is presented.
Approach: They propose a new dataset for classifying memes by their template and communicative intent.
Outcome: The proposed method outperforms existing methods in classifying memes by their template and communicative intent.
No, you’re not alone: A better way to find people with similar experiences on Reddit (D19-55)

Copied to clipboard

Challenge: a probabilistic clustering algorithm can help users find posts that discuss experiences similar to their own . a recent study shows that probabilistic Clustering can yield a better performance than baseline clustering methods .
Approach: They propose a probabilistic clustering algorithm that can help Reddit users find posts that discuss experiences similar to their own.
Outcome: The proposed algorithm can find posts that discuss experiences similar to their own . it performs better than baseline clustering methods due to high runtime overhead .
Modeling Ideological Salience and Framing in Polarized Online Groups with Graph Neural Networks and Structured Sparsity (2022.findings-naacl)

Copied to clipboard

Challenge: Existing methods to detect ideological divides in social media rely on knowing in advance the political orientation of text . fascist and mainstream are among the most polarized concepts in reddit in 2019 .
Approach: They propose a minimally supervised method that leverages the network structure of online discussion forums to detect polarized concepts.
Outcome: The proposed framework captures temporal ideological dynamics such as right-wing and left-wing radicalization using graph neural networks and sparsity learning.
A System for Dynamically Tracking Content Moderation on Reddit (2026.acl-demo)

Copied to clipboard

Challenge: Recent work in social media platforms delegate content moderation decisions to users and communities.
Approach: They propose a software system for the dynamic monitoring of Reddit posts, communities, and moderation actions to enable scalable and reproducible research on decentralized platform governance and content moderation.
Outcome: The proposed system is the only available solution for general-purpose, real-time, policy-compliant longitudinal data collection on Reddit.
A Word is Worth A Thousand Dollars: Adversarial Attack on Tweets Fools Stock Prediction (2022.naacl-main)

Copied to clipboard

Challenge: Existing models are vulnerable to adversarial attacks, but their vulnerability is underexplored.
Approach: They propose to concatenate a perturbed but semantically similar tweet into a model that fools stock prediction models.
Outcome: The proposed method achieves consistent success rates and causes significant monetary loss in trading simulation by simply concatenating a perturbed but semantically similar tweet.
Classification without (Proper) Representation: Political Heterogeneity in Social Media and Its Implications for Classification and Behavioral Analysis (2022.findings-acl)

Copied to clipboard

Challenge: Prior work has shown that partisan leanings can be inferred from a diverse set of behavioral characteristics such as text, social networks, and even community participation.
Approach: They test this assumption and show that commonly-used models do not generalize . they also show that political users are more toxic on the platform and inter-party interactions are even more toxic .
Outcome: The proposed models do not generalize, indicating heterogeneous political users.
MTNT: A Testbed for Machine Translation of Noisy Text (D18-1)

Copied to clipboard

Challenge: Noisy input text can cause disastrous mistranslations in most modern machine translation systems.
Approach: They propose a benchmark dataset for Machine Translation of Noisy Text (MTNT) they use reddit comments and professionally sourced translations to examine noise types.
Outcome: The proposed dataset can provide an attractive testbed for noise-robust machine translation systems.
Breaking Down the Invisible Wall of Informal Fallacies in Online Discussions (2021.acl-long)

Copied to clipboard

Challenge: a number of people engage in unsound argumentation techniques to prove a claim on online platforms . fallacies are weak arguments that seem convincing, but their evidence does not prove or disprove the conclusion .
Approach: They propose to use user comments containing fallacy mentions as noisy labels to classify fallacies . they use the pragma-dialectical theory of argumentation to study the most common fallacias on Reddit .
Outcome: The proposed dataset of fallacies on reddit shows that neural models perform better in conversational context.
Subjective Crowd Disagreements for Subjective Data: Uncovering Meaningful CrowdOpinion with Population-level Learning (2023.acl-long)

Copied to clipboard

Challenge: Annotator disagreements are resolved before learning takes place, but researchers question the performance of a system when annotators disagree.
Approach: They propose a method that uses language features and label distributions to pool similar items into larger labels.
Outcome: The proposed method is based on five publicly available datasets with varying levels of disagreements on social media and in the wild using a dataset from Facebook.
Analyzing Norm Violations in Live-Stream Chat (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for detecting toxic language and norm violations are limited to live-streaming platforms . existing methods are less effective when applied to live streaming platforms based on a limited time frame .
Approach: They propose to use contextual information to automatically moderate toxic content on live streaming platforms.
Outcome: The proposed model improves on live-streaming platforms by 35%.
Sentence-Level Content Planning and Style Specification for Neural Text Generation (D19-1)

Copied to clipboard

Challenge: Recent advances in text generation systems often produce incoherent and unfaithful outputs . a novel automated text generation system takes into account content selection, text planning, and surface realization.
Approach: They propose an end-to-end trained two-step text generation model that considers sentence-level content planners and language styles.
Outcome: The proposed model outperforms competing models in three domains with diverse topics and varying language styles.
RedHOT: A Corpus of Annotated Medical Questions, Experiences, and Claims on Social Media (2023.findings-eacl)

Copied to clipboard

Challenge: Social media platforms such as Reddit are vulnerable to misinformation and disinformation.
Approach: They propose a method to automatically derive (noisy) supervision for retrieval of trustworthy evidence relevant to a given claim made on social media.
Outcome: The proposed method outperforms baseline models in the retrieval task performed by medical doctors.
The Change that Matters in Discourse Parsing: Estimating the Impact of Domain Shift on Parser Error (2022.findings-acl)

Copied to clipboard

Challenge: Discourse analysis is very low on texts outside of the training distribution’s coverage, diminishing the practical utility of existing models.
Approach: They propose to use a distribution shift statistic to estimate the error-gap of a discourse model and to use it to estimate it.
Outcome: The proposed model can be estimated via distribution shift but does not correlate with change in the observed error of a classifier (i.e. error-gap).
DS-TOD: Efficient Domain Specialization for Task-Oriented Dialog (2022.findings-acl)

Copied to clipboard

Challenge: Recent work shows that self-supervised dialog-specific pretraining on large conversational datasets yields substantial gains over traditional language modeling (LM) pretraining.
Approach: They propose a resource-efficient and modular domain specialization by means of domain adapters in which domain knowledge is encoded.
Outcome: The proposed framework extracts domain-specific terms and then uses them to build DomainCC and DomainReddit resources based on masked language modeling and response selection objectives.
QuaLLM: An LLM-based Framework to Extract Quantitative Insights from Online Forums (2025.findings-naacl)

Copied to clipboard

Challenge: Qualitative and quantitative methods to analyze text data on online forums are infeasible to scale or require significant human effort to translate outputs to human readable forms.
Approach: They propose a novel LLM-based framework to analyze and extract quantitative insights from text data on online forums.
Outcome: The proposed framework analyzes over one million comments from two of Reddit’s rideshare worker communities, marking the largest study of its type.
Investigating Human Values in Online Communities (2025.naacl-long)

Copied to clipboard

Challenge: Existing value frameworks struggle with sample sizes and rely on selfreported surveys to calculate values.
Approach: They propose a method to computationally analyse values on Reddit using in-domain and out-of-domain human annotations to train a value relevance and a polarity classifier.
Outcome: The proposed method can be used to analyse values on reddit using human annotations and human annotation.
Reasoning with Sarcasm by Reading In-Between (P18-1)

Copied to clipboard

Challenge: Sarcasm is a figurative speech act which manifests on social networks such as Twitter and Reddit.
Approach: They propose a model that looks in-between rather than across to explicitly model contrast and incongruity.
Outcome: The proposed model achieves state-of-the-art performance on all datasets and improves interpretability.
PairScale: Analyzing Attitude Change with Pairwise Comparisons (2025.findings-naacl)

Copied to clipboard

Challenge: a text-based framework for measuring attitudes in communities is proposed . the framework uses both implicit and explicit evidence in language to characterize attitudes .
Approach: They propose a text-based framework for measuring attitudes in communities toward issues of interest using language.
Outcome: The proposed framework is validated by examining attitudes on two high-profile issues in the u.s.
LLMs Cannot (Yet) Match the Specificity and Simplicity of Online Communities in Long Form Question Answering (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent years have positioned Large Language Models (LLMs) as powerful question answering (QA) tools, shifting users away from interacting in communities towards discourse with AI-driven conversational interfaces.
Approach: They propose to use a QA preference dataset to fine-tune and align Large Language Models (LLMs) from more than 7.4 million submissions and 82 million comments from 2008 to 2022 in Reddit’s 15 largest finance communities.
Outcome: The proposed framework improves on the social quality of the data, and the proposed framework is more accurate and more specific.
Digital Gatekeepers: Google’s Role in Curating Hashtags and Subreddits (2025.acl-long)

Copied to clipboard

Challenge: This study examines how search engines like Google selectively promote or suppress certain hashtags and subreddits, impacting the flow of information and impacting public conversations.
Approach: They compare search engine results with nonsampled data from Reddit and Twitter/X to examine how search engines curate content through algorithmic curation.
Outcome: The proposed algorithm suppresses subreddits related to sexually explicit material, conspiracy theories, advertisements, and cryptocurrencies while promoting content associated with higher engagement.
Modeling Intra and Inter-modality Incongruity for Multi-Modal Sarcasm Detection (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for sarcasm detection ignore the incongruity character in sarcasm, which is often manifested between modalities or within modalités.
Approach: They propose to capture inter-modality incongruity in a text-based model by using a self-attention mechanism and a co-attention model to model the contradiction within the text.
Outcome: The proposed model achieves state-of-the-art on a public multi-modal sarcasm detection dataset.
Weakly-Supervised Methods for Suicide Risk Assessment: Role of Related Domains (2021.acl-short)

Copied to clipboard

Challenge: Among social media platforms, Reddit has emerged as the most promising one due to its anonymity and its focus on topic-based communities (subreddits) . a challenge for previous work on suicide risk assessment has been the small amount of labeled data.
Approach: They propose to use social media to collect user data from r/SuicideWatch subreddit and annotate it with user-level suicide risk: no-risk, low-risk and high-risk.
Outcome: The proposed model improves by using pseudo-labeling based on related issues around mental health (e.g., anxiety, depression)
CLICK: Contrastive Learning for Injecting Contextual Knowledge to Conversational Recommender System (2023.eacl-main)

Copied to clipboard

Challenge: Existing CRSs lack capturing comprehensive user preferences . existing systems lack contextual knowledge to capture user preferences from a dialogue context .
Approach: They propose a Contrastive Learning approach for Injecting Contextual Knowledge from Reddit data to a CRS task.
Outcome: The proposed approach captures a user preference from a dialogue context without items . it improves on the existing methods, and the results are published in the journal of cognitive science.
RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language Models (2021.acl-long)

Copied to clipboard

Challenge: Recent work has focused on measuring and mitigating bias in pretrained language models.
Approach: They propose a dataset that measures and mitigates bias across gender,race, religion, and queerness . they compare REDDITBIAS to a widely used conversational DialoGPT model .
Outcome: The proposed framework measures and mitigates bias across gender,race, religion, and queerness dimensions.
What Sounds “Right” to Me? Experiential Factors in the Perception of Political Ideology (2021.eacl-main)

Copied to clipboard

Challenge: a recent study has suggested that political ideology is inherently built into text . a new study examines the impact of experiential factors on annotator perceptions of political ideology .
Approach: They propose to investigate the impact of experiential factors on annotator perceptions of political ideology by analyzing an annotated corpus of political discussion in the U.S. They find that these factors may influence consistency of how political ideologies are perceived by annotators.
Outcome: The findings challenge the assumption that political ideology is built into text . they show that experiential factors may influence how ideologies are perceived .
Personalized Response Generation via Generative Split Memory Network (2021.naacl-main)

Copied to clipboard

Challenge: Despite the success of text generation and dialogue systems, how to endow a text generation system with personality traits remains under-investigated.
Approach: They propose a model to generate personalized responses on reddit using user profiles and posting histories.
Outcome: The proposed model improves over the state-of-the-art response generation models.
Identifying Self-Disclosures of Use, Misuse and Addiction in Community-based Social Media Posts (2024.findings-naacl)

Copied to clipboard

Challenge: Experimental results show that identifying the phases of opioid use disorder is highly contextual and challenging.
Approach: They analyze 2500 opioid-related posts from various subreddits labeled with six different phases of opioid use . they annotate span-level extractive explanations and critically evaluate state-of-the-art models in a supervised, few-shot, or zero-shot setting.
Outcome: The proposed models improve classification accuracy and quality of the extracted explanations.
Open-Mindedness and Style Coordination in Argumentative Discussions (2021.eacl-main)

Copied to clipboard

Challenge: Previous research has shown that linguistic accommodation correlates with gaps in the power and status of the speakers and the way it promotes approval and discussion efficiency.
Approach: They propose a novel perspective on linguistic accommodation, exploring its correlation with the open-mindedness of a speaker, rather than to her social status.
Outcome: The proposed approach improves the open-mindedness of a speaker and lowers discussion efficiency.
Learning to Disentangle Interleaved Conversational Threads with a Siamese Hierarchical Network and Similarity Ranking (N18-1)

Copied to clipboard

Challenge: Existing methods to disentangle interleaved conversations can lead to difficulties in following discussions and retrieving relevant information from simultaneous messages.
Approach: They propose to leverage representation learning to separate intermingled messages into detached conversations by estimating conversation-level similarity between closely posted messages.
Outcome: The proposed approach outperforms baselines in pairwise similarity estimation and conversation disentanglement.
Social Meme-ing: Measuring Linguistic Variation in Memes (2024.naacl-long)

Copied to clipboard

Challenge: In this paper, we analyze memes as a form of language subject to the same kinds of sociolinguistic variation as other modalities, such as written language and speech.
Approach: They propose a computational pipeline to cluster memes into templates and semantic variables, taking advantage of their multimodal structure to learn meme semantics from an unstructured dataset.
Outcome: The proposed method uses 3.8M images from a reddit meme database to analyze linguistic variation in memes.
Learning Invariant Representations of Social Media Users (D19-1)

Copied to clipboard

Challenge: Existing methods for learning to compare social media users fail to generalize to new users or even to previously known users.
Approach: They propose a procedure to learn a mapping from short episodes of user activity to a vector space in which the distance between points captures the similarity of the corresponding users’ invariant features.
Outcome: The proposed procedure can be applied to users not seen at training time and enables efficient comparisons of users in the resulting vector space.
Re-entry Prediction for Online Conversations via Self-Supervised Learning (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing work on re-entry prediction ignores conversation thread patterns and repeated engagement of target users.
Approach: They propose to use conversation thread patterns to predict whether a user will come back to a conversation they once participated in to train a model on labels that are automatically derived from the data.
Outcome: The proposed task outperforms the state-of-the-art models on two social media datasets with fewer parameters and faster convergence.
DPL: Diverse Preference Learning Without A Reference Model (2025.naacl-long)

Copied to clipboard

Challenge: Existing methods to direct preference alignment do not utilize diversity in preference annotations which limits their applicability.
Approach: They propose a reference-model-free method that learns a baseline desirability in LLM responses while being robust to the diversity of preference annotations.
Outcome: The proposed method learns a baseline desirability in LLM responses while being robust to the diversity of preference annotations.
The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies have examined the impact of recommendation algorithms on how users discover and join online groups, but there are few standardized datasets for generating such models.
Approach: They propose to use Reddit to build a dataset that can be used to build models of user engagement with online groups.
Outcome: The proposed model is based on the behavior of subreddits banned in June 2020 as part of Reddit's efforts to stop the dissemination of hate speech.
Plug-and-Play Conversational Models (2020.findings-emnlp)

Copied to clipboard

Challenge: Large conversational models that generate coherent and fluent responses often require large dialogue datasets.
Approach: They propose and evaluate plug-and-play methods for controllable response generation . they demonstrate a high degree of control over the generated conversational responses .
Outcome: The proposed method does not require further computation at decoding time and does not need fine-tuning of a large language model.
A Comprehensive Study of Gender Bias in Chemical Named Entity Recognition Models (2024.naacl-long)

Copied to clipboard

Challenge: Chemical named entity recognition (NER) models are used in many downstream tasks, but it is unknown whether they work the same for everyone.
Approach: They develop a framework for measuring gender bias in chemical NER models . they analyze a corpus of 92,405 words with self-identified gender information from reddit .
Outcome: The proposed framework measures gender bias in chemical NER models using synthetic data and a newly annotated corpus of over 92,405 words with self-identified gender information from Reddit.
Argument Generation with Retrieval, Planning, and Realization (P19-1)

Copied to clipboard

Challenge: a novel argument generation framework is used to generate counter-arguments . CANDELA uses a text planning decoder to retrieve arguments of different perspectives .
Approach: They propose a powerful retrieval system and a novel two-step argument generation framework . they use a retrieval-based retrieval platform indexed with 12 million articles from Wikipedia .
Outcome: The proposed framework yields higher BLEU, ROUGE, and METEOR scores than state-of-the-art models.
DDisCo: A Discourse Coherence Dataset for Danish (2022.lrec-1)

Copied to clipboard

Challenge: Discourse coherence models have been developed using randomly shuffled texts instead of highly edited and coherent data.
Approach: They propose to annotate Danish Wikipedia and Reddit for discourse coherence using real-world text instead of artificially incoherent text for training and testing models.
Outcome: The proposed model performs well on annotated texts from the Danish Wikipedia and Reddit dataset.
Joint Effects of Context and User History for Predicting Online Conversation Re-entries (P19-1)

Copied to clipboard

Challenge: Existing methods for predicting online conversation re-entry focus on modeling engagement patterns in ongoing conversations or ignoring the rich information in users' previous chatting history.
Approach: They propose a neural framework with three main layers to model the conversation context and user history and their interactions with Twitter and Reddit to predict whether a user will return to a conversation they once participated in.
Outcome: The proposed framework outperforms the state-of-the-art methods on two large-scale Twitter and Reddit conversations, and achieves an F1 score of 61.1 on Twitter conversations.
Echoes of Discord: Forecasting Hater Reactions to Counterspeech (2025.findings-naacl)

Copied to clipboard

Challenge: Hate speech (HS) online causes increased prejudice and discrimination, fostering an environment of hostility and social division.
Approach: They analyze the Reddit Echoes of Hate dataset to assess the impact of counterspeech from the hater's perspective and focus on whether the counterspeak leads the reentry to be hateful.
Outcome: The proposed model outperforms the two-stage reaction predictor and the three-way classifier to predict haters' reactions to the reentry of the conversation and determines the type of resentment.
MentSum: A Resource for Exploring Summarization of Mental Health Online Posts (2022.lrec-1)

Copied to clipboard

Challenge: Mental health remains a significant challenge of public health worldwide . many use online platforms to share their mental health conditions and seek help .
Approach: They analyze a dataset of over 24k user posts from Reddit and 43 mental health subreddits to generate a short summarization.
Outcome: The proposed dataset compared over 24k user posts and 43 mental health subreddits . it shows that the summarization of these posts is faster and more accurate than previous studies.
CultureBank: An Online Community-Driven Knowledge Base Towards Culturally Aware Language Technologies (2024.findings-emnlp)

Copied to clipboard

Challenge: CultureBank is a knowledge base built upon users’ self-narratives with 12K cultural descriptors sourced from TikTok and 11K from Reddit.
Approach: They construct a pipeline to construct cultural knowledge bases from different online communities on a massive scale.
Outcome: The proposed pipeline improves cultural awareness of language models by evaluating them on two cultural tasks in a zero-shot setting.
Curating a Large-Scale Motivational Interviewing Dataset Using Peer Support Forums (2022.coling-1)

Copied to clipboard

Challenge: Existing therapeutic chatbots lack large-scale conversations between clients and trained counselors . prior work has found that social media platforms such as Reddit are used to vent distress and peers are seen to actively respond to such posts.
Approach: They propose to use peer support platforms to scrape conversational data from Reddit to determine whether counselors' responses align with real therapeutic conversations.
Outcome: The proposed method achieved 97% coverage out of 17.3K responses, meaning that out of 16.8K responses labeled with a moderate agreement.
Automatic Generation of Large-scale Multi-turn Dialogues from Reddit (2022.coling-1)

Copied to clipboard

Challenge: Using a set of algorithms, we can generate large dialogue corpus from Reddit.
Approach: They propose to automatically convert posts and their comments from discussion forums such as Reddit into multi-turn dialogues.
Outcome: The proposed methods improve on the baseline method by 36.3% . the best method shows an improvement of 36.6% over the previous one .
Using RL to Identify Divisive Perspectives Improves LLMs Abilities to Identify Communities on Social Media (2024.findings-emnlp)

Copied to clipboard

Challenge: Experimental results show improvements on Reddit and Twitter data .
Approach: They propose to take advantage of Large Language Models (LLMs) to better identify user communities.
Outcome: The proposed model improves on Reddit and Twitter data and tasks of community detection, bot detection, and news media profiling.
From the Stage to the Audience: Propaganda on Reddit (2021.eacl-main)

Copied to clipboard

Challenge: a recent opinion piece in the Washington Post highlights a difference between the political discourse in the two countries.
Approach: They analyze political forums on Reddit that target a diverse audience in two countries . they focus on three research questions: who is posting propaganda? and how is propaganda received?
Outcome: The authors analyze political forums on reddit in the US and the UK for one year . they find that propaganda is misleading and is received by different audiences .
Integrating Tree Structures and Graph Structures with Neural Networks to Classify Discussion Discourse Acts (C18-1)

Copied to clipboard

Challenge: Existing models that analyze textual contents and discussion structures require understanding of textual content and discussion structure.
Approach: They propose a model that integrates discussion structures with neural networks to classify discourse acts.
Outcome: The proposed model improves accuracy and FB1 score by 1.5% compared to the previous best model.
DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation (2022.acl-long)

Copied to clipboard

Challenge: Existing pre-trained dialog models shed light on various downstream tasks in natural language processing (NLP).
Approach: They propose a dialog pre-training framework that introduces latent variables into the enhanced encoder-decoder pre-train framework to increase relevance and diversity of responses.
Outcome: The proposed model achieves state-of-the-art on personaChat, DailyDialog, and DSTC7-AVSD datasets.
FACTOID: A New Dataset for Identifying Misinformation Spreaders and Political Bias (2022.lrec-1)

Copied to clipboard

Challenge: Proactively identifying misinformation spreaders is an important step towards mitigating the impact of fake news on our society.
Approach: They propose a new reddit dataset for fake news spreader analysis, called FACTOID, which tracks political discussions on Reddit since the beginning of 2020.
Outcome: The proposed dataset contains over 4K users with 3.4M posts and includes their credibility level (very low to very high) and political bias strength (extreme right to extreme left).
Leveraging Training Dynamics and Self-Training for Text Classification (2022.findings-emnlp)

Copied to clipboard

Challenge: Semi-supervised learning (SSL) is a promising technique for improving deep learning models when training data is scarce.
Approach: They propose a semi-supervised learning approach that leverages training dynamics of unlabeled data.
Outcome: The proposed method achieves an average increase in F1 score of 3.5% over baselines in low resource settings.
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing (2026.eacl-long)

Copied to clipboard

Challenge: a single prompt can inspire countless valid stories, making objective verification impossible.
Approach: They propose a large-scale benchmark for creative writing evaluation using a reddit corpus and a 2,480-pair test set.
Outcome: The proposed model outperforms existing OTS judges and generative reward models in the evaluation of creative writing.
ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations (2026.eacl-long)

Copied to clipboard

Challenge: Evaluating 9 state-of-the-art LLMs reveals two critical limitations: 61% of incorrect span predictions are semantically unrelated to actual errors.
Approach: They propose a benchmark of 1,001 expert-annotated question-answer pairs with span-level error annotations derived from Reddit's r/AskScience.
Outcome: Evaluating 9 state-of-the-art LLMs, we find that comparative judgment is paradoxically harder than independent detection when comparing answers side-by-side.
Hope ‘The Paragraph Guy’ explains the rest : Introducing MeSum, the Meme Summarizer (2024.findings-emnlp)

Copied to clipboard

Challenge: a lack of large datasets for supervised learning and resource-intensive vision language models have hindered the development of meme comprehension.
Approach: They propose a framework to bridge the gap between meme comprehension and vision language models by using a multimodal dataset.
Outcome: The proposed framework outperforms existing methods in the meme comprehension test.
Aligning Multidimensional Worldviews and Discovering Ideological Differences (2021.emnlp-main)

Copied to clipboard

Challenge: Existing work on understanding worldviews and ideological distinctions focuses on political polarization . et al., 2018: a novel method for uncovering complex ideological and worldview characteristics of communities.
Approach: They propose a method to uncover multifaceted ideological differences across multiple axes . they use comments from the largest communities on reddit.com to train word embedding models .
Outcome: The proposed method can uncover complex ideological differences across multiple axes of polarization using over 1B comments from the largest communities on reddit.com representing 40% of Reddit activity.
Cryptocurrency Bubble Detection: A New Stock Market Dataset, Financial Task & Hyperbolic Models (2022.naacl-main)

Copied to clipboard

Challenge: speculative trading of highly volatile assets such as cryptocurrencies and meme stocks presents a new challenge in the financial realm.
Approach: They propose a multi-span bubble detection task based on social media hype and a set of sequence-to-sequence hyperbolic models . they use data from 9 exchanges over five years to test their models based upon the power-law dynamics of cryptocurrencies and user behavior on social networks.
Outcome: The proposed model is able to detect bubbles on a set of reddit and twitter posts spanning over two million tweets over five years .
ComPO: Community Preferences for Language Model Personalization (2025.naacl-long)

Copied to clipboard

Challenge: Current methods for training language models with human feedback rely on subjective preferences that are assumed to account for an "average" user . however, annotating preferences is inherently subjective and results in generic models that generate outputs not preferred by many user groups.
Approach: They propose a method to personalize preference optimization in LMs by contextualizing the probability distribution of model outputs with the preference provider.
Outcome: The proposed method improves performance by focusing on group-level preferences rather than individual feedback.
Help! Need Advice on Identifying Advice (2020.emnlp-main)

Copied to clipboard

Challenge: Pre-trained systems are able to capture advice better than rule-based systems, but advice identification is challenging.
Approach: They analyze a dataset of advice posts on two reddit forums and annotate whether they contain advice.
Outcome: The proposed models show that pre-trained models capture advice better than rule-based systems, but advice identification is challenging.
CHARM: Inferring Personal Attributes from Conversations (2020.emnlp-main)

Copied to clipboard

Challenge: Personal Knowledge Bases (PKBs) capture individual user traits for customizing downstream applications like chatbots or recommenders.
Approach: They propose a method that leverages keyword extraction and document retrieval to predict attribute values that were never seen during training.
Outcome: The proposed method can predict attributes that were never seen during training.
REALM: A Dataset of Real-World LLM Use Cases (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies on LLM adoption and their social implications lack empirical grounding, weakening their validity.
Approach: They propose to integrate a dataset of over 94,000 LLM use cases collected from Reddit and news articles to provide insights into LLM adoption across different domains.
Outcome: The proposed dataset includes over 94,000 LLM use cases collected from Reddit and news articles.
Making “fetch” happen: The influence of social and linguistic context on nonstandard word growth and decline (D18-1)

Copied to clipboard

Challenge: In an online community, new words come and go, but language change is shaped and constrained by the grammatical system in which it takes part.
Approach: They analysed the frequency of non-standard words in reddit to determine their impact on language change.
Outcome: The results show that language change is shaped and constrained by the grammatical system in which it takes place.
Evaluating Verifiability in Generative Search Engines (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing generative search engines are rapidly gaining users, according to a new study . existing systems are poorly cited and lack reliability, a study finds .
Approach: They conduct human evaluations of four popular generative search engines . they find that existing generative engines are fluent and appear informative .
Outcome: The results show that existing generative search engines are not reliable and often contain unsupported statements and inaccurate citations.
Beyond Discrete Personas: Personality Modeling Through Journal Intensive Conversations (2025.coling-main)

Copied to clipboard

Challenge: Existing LLMs rely on static, predefined personas to capture dynamic and evolving nature of human personalities.
Approach: They propose a dataset with 400,000 conversations and a framework for generating personalized conversations using long-form journal entries from Reddit.
Outcome: The proposed framework generates high-quality, personality-rich dialogues grounded in reddit journal entries.
Neural Conversation Recommendation with Online Interaction Modeling (D19-1)

Copied to clipboard

Challenge: Existing models that only use lexical features and ignore past user interactions in online conversations are inadequate to identify and engage in online discussions.
Approach: They propose a framework that automatically recommends conversations based on user's prior conversation behaviors by exploring deep semantic features that measure how a user’s preferences match an ongoing conversation’s context.
Outcome: The proposed model outperforms state-of-the-art models on two large-scale datasets from Twitter and Reddit showing that it incorporates deep semantic features that measure how a user’s preferences match an ongoing conversation’s context.
A Benchmark Dataset for Learning to Intervene in Online Hate Speech (D19-1)

Copied to clipboard

Challenge: Existing methods to detect online hate speech ignore conversational context . generative hate speech intervention is a novel approach to counter online hate .
Approach: They propose a task where generative hate speech intervention generates responses to intervene during online conversations that contain hate speech.
Outcome: The proposed method can detect and block hate speech and discourage it . it can also generate responses written by Mechanical Turk workers .
Words Matter: Reducing Stigma in Online Conversations about Substance Use with Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Only 7% of people living with an SUD receive any form of treatment, with stigma reported as a major barrier.
Approach: They propose a computational framework for analyzing stigma and de-stigmatizing online content and delving into the linguistic features that propagate stigma towards PWUS.
Outcome: The proposed model transforms stigmatizing language into more empathetic language and analyzes over 1.2 million posts on social media .
Generating similes effortlessly like a Pro: A Style Transfer Approach for Simile Generation (2020.emnlp-main)

Copied to clipboard

Challenge: Literary tropes are at the crux of human imagination and communication.
Approach: They propose to automatically transform similes from reddit to their literal counterparts using common sense knowledge to generate simile models.
Outcome: The proposed method generates 88% novel similes that do not share properties with training data.
APPDIA: A Discourse-aware Transformer-based Style Transfer Model for Offensive Social Media Conversations (2022.coling-1)

Copied to clipboard

Challenge: Using style-transfer models to reduce offensiveness of social media comments is difficult because of limited labeled data.
Approach: They propose two methods to integrate discourse relations with pretrained style-transfer models and evaluate them on a reddit dataset.
Outcome: The proposed models can reduce offensiveness while preserving original meaning . they are the first to examine inferential links between comment and original text .
Like a Therapist, But Not: Reddit Narratives of AI in Mental Health Contexts (2026.findings-acl)

Copied to clipboard

Challenge: Large language models are increasingly used for emotional support and mental health–related interactions outside clinical settings.
Approach: They analyze 5,126 Reddit posts describing use of AI for emotional support or therapy . positive sentiment is most strongly associated with task and goal alignment, they say .
Outcome: The proposed framework analyzes language, adoption-related attitudes, and relational alignment at scale. positive sentiment is most strongly associated with task and goal alignment.
Enhancing Discourse Parsing for Local Structures from Social Media with LLM-Generated Data (2025.coling-main)

Copied to clipboard

Challenge: Existing discourse parsers do not generalize well across genres and text types.
Approach: They propose to integrate large language models into RST discourse parsers to improve parser performance in a social media context.
Outcome: The proposed model improves parser performance in a social media context without pre-identified discourse units.
Among Us: Language of Conspiracy Theorists on Mainstream Reddit (2026.acl-long)

Copied to clipboard

Challenge: Conspiracy theories are influential, alternative narratives that explain events through the actions of secretive, malevolent groups.
Approach: They analyze a large-scale longitudinal dataset of over 500 million comments on reddit . they show that users exhibit distinctive linguistic patterns that enable machine learning models to distinguish them from the general population within individual communities.
Outcome: The proposed model outperforms global classifiers by 17 percentage points.
EmoMent: An Emotion Annotated Mental Health Corpus from Two South Asian Countries (2022.coling-1)

Copied to clipboard

Challenge: Recent research using AI and NLP demonstrates strong potential to automatically detect mental health issues from digital footprints such that professionals could provide timely interventions and mental health resources to vulnerable persons.
Approach: They developed an emotion-annotated mental health corpus from 2802 Facebook posts extracted from two South Asian countries, Sri Lanka and India.
Outcome: The proposed model achieved 98.3% agreement between the annotators and a Fleiss’ Kappa of 0.82.
SGCD: Subtask-Guided Causal-Debiasing Framework for Robust Cross-Utterance Sentiment Quadruple Extraction in Dialogues (2025.findings-emnlp)

Copied to clipboard

Challenge: a new framework for sentiment analysis in dialogues addresses cross-utterance elements and focus biases . SGCD framework employs multi-granularity attention paths to enhance cross-interaction matching .
Approach: a framework is developed to help analyze sentiments in multi-turn dialogues . it leverages subtask-specific features to guide learning of token-level features .
Outcome: The proposed framework outperforms state-of-the-art methods in analyzing conversational data . cross-utterance elements and focus bias are challenges, authors say .
TalkUp: Paving the Way for Understanding Empowering Language (2023.findings-emnlp)

Copied to clipboard

Challenge: Empowerment has rarely been studied in NLP because of its implicit nature . linguistics and psychology research shows how empowerment can impact people by increasing their sense of self-efficacy and self-esteem.
Approach: They crowdsource Reddit posts labeled for empowerment and use it to train language models that capture empowering and disempowering language.
Outcome: The proposed dataset can be used to train language models that capture empowering and disempowering language.
Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and Analysis (2022.emnlp-main)

Copied to clipboard

Challenge: Prior work on ideology prediction has focused on single modalities, i.e., text or images.
Approach: They propose a task where a model predicts binary or five-point scale ideological leanings given a text-image pair with political content.
Outcome: The proposed model outperforms the state-of-the-art model by almost 4% and a strong multimodal baseline with no pretraining by over 3%.
Answering Narrative-Driven Recommendation Queries via a Retrieve–Rank Paradigm and the OCG-Agent (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to generate narrative-driven recommendation are based on large language models (LLMs) but the RAG paradigm is inherently ill-suited for such special queries.
Approach: They propose a novel retrieve-rank paradigm that generatively retrieves structurally adaptive and semantically aligned candidates, ensuring both extensive candidate coverage and high-quality information.
Outcome: The proposed paradigm outperforms the existing paradigm and the existing one under real-world scenarios.
SPADE: A Big Five-Mturk Dataset of Argumentative Speech Enriched with Socio-Demographics for Personality Detection (2022.lrec-1)

Copied to clipboard

Challenge: Recent efforts to create such datasets from social media do not include continuous and contextualized language use.
Approach: They propose to use argumentative speech to generate a dataset with continuous arguments labeled with the Big Five personality traits and enriched with socio-demographic data.
Outcome: The proposed model leverages 436 (psycho)linguistic features extracted from transcribed speech and speaker-level metainformation with transformers to investigate which types of features contribute to the prediction of individual personality traits.
NeuTral Rewriter: A Rule-Based and Neural Approach to Automatic Rewriting into Gender Neutral Alternatives (2021.emnlp-main)

Copied to clipboard

Challenge: Recent years have seen an increasing need for gender-neutral and inclusive language.
Approach: They propose a rule-based and a neural approach to gender-neutral rewriting for English . they use manually curated synthetic and natural data to train a rewriter .
Outcome: The proposed approach improves on the rule-based approach with word error rates below 0.18% on synthetic, in-domain and out-domain test sets.
CryptOpiQA: A new Opinion and Question Answering dataset on Cryptocurrency (2025.coling-main)

Copied to clipboard

Challenge: Using a dataset of tweets and Reddit, we investigate the public opinion on cryptocurrency and bitcoin on Twitter and RedDit.
Approach: They create a dataset to investigate the public opinion on cryptocurrency and bitcoin on Twitter and Reddit.
Outcome: The proposed dataset contains gold standard and silver standard labels and a question-answering sub-corpus.
A Corpus of German Reddit Exchanges (GeRedE) (2020.lrec-1)

Copied to clipboard

Challenge: Reddit is a popular online platform combining social news aggregation, discussion and microblogging.
Approach: They propose a method to filter out German data and further pre-processing steps to find out what is linguistically peculiar in the German data.
Outcome: The proposed method filters out German data and includes metadata and annotation layers.
Using Sociolinguistic Variables to Reveal Changing Attitudes Towards Sexuality and Gender (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies show that word choice is driven by demographics within the United States.
Approach: They develop computational methods to study word choice within a sociolinguistic lexical variable . they use two variables to test for attitudes towards sexuality and gender in the u.s.
Outcome: The proposed methods allow us to examine attitudes towards sexuality and gender in the United States through two lexical variables.
Misery Loves Complexity: Exploring Linguistic Complexity in the Context of Emotion Detection (2023.findings-emnlp)

Copied to clipboard

Challenge: a negative emotion is a cognitive bias that affects how we express thoughts and opinions online . a recent study shows that negative words generate more engagement and clicks than positive ones .
Approach: They propose to use readability and linguistic complexity metrics to better understand emotions . they propose to fine-tune three state-of-the-art transformers to detect emotions based on a dataset .
Outcome: The proposed model fails to predict emotions on complex texts, the authors show . they also show that more advanced models fail to predict complex texts .
LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming (2023.acl-long)

Copied to clipboard

Challenge: a recent study shows that open-domain dialogue systems are not able to perform well in fast-growing scenarios such as live streaming due to the domain gap between online-post constructed data and those required in downstream conversational tasks.
Approach: They propose to train a conversational agent based on large social media datasets with multiple domains to improve response in live streaming scenarios.
Outcome: The proposed model improves response modeling and addressee recognition in live open-domain scenarios.
Community-Cross-Instruct: Unsupervised Instruction Generation for Aligning Large Language Models to Online Communities (2024.emnlp-main)

Copied to clipboard

Challenge: Social scientists use surveys to learn opinions and beliefs of populations, but these methods are slow, costly, and prone to biases.
Approach: They propose a framework for aligning large language models to online communities by finetuning instruction-output pairs by an advanced LLM to elicit their beliefs.
Outcome: The proposed framework enables cost-effective and automated surveying of diverse online communities.
Measuring Psychological Depth in Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Current evaluations of creative stories focus on objective properties of the text, such as its style, coherence, diversity, and creativity.
Approach: They propose a framework that measures an LLM's ability to produce authentic and narratively complex stories that provoke emotion, empathy, and engagement.
Outcome: The proposed framework shows that humans can consistently evaluate stories based on the PDS (0.72 Krippendorff’s alpha).
MentalHelp: A Multi-Task Dataset for Mental Health in Social Media (2024.lrec-main)

Copied to clipboard

Challenge: Annotating social media data for mental health disorders is expensive and time-consuming, limiting their size and scope.
Approach: They present a large-scale semi-supervised mental disorder detection dataset containing 14 million instances from Reddit and an ensemble of three separate models.
Outcome: The proposed dataset contains 14 million instances of mental disorders . it was collected from reddit and labeled in a semi-supervised way .
Parallel Communities Across the Surface Web and the Dark Web (2025.findings-emnlp)

Copied to clipboard

Challenge: Sense of Community is a social motivation that is reflected in the social behavior of humans.
Approach: They compile a large collection of parallel community datasets comprising over 7 million posts and comments from Reddit and 200,000 posts and comment from Dread, a dark web discussion forum, covering similar topics.
Outcome: The results show that users on Reddit exhibit a stronger sense of community membership despite the dark web’s restricted accessibility.
Can Large Language Models Identify Implicit Suicidal Ideation? An Empirical Evaluation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing data on suicidal ideation in private conversations are limited . a new dataset of 1,200 test cases is presented to address this gap .
Approach: They propose a dataset of 1,200 test cases simulating implicit suicidal ideation in private contexts.
Outcome: The proposed dataset includes 1,200 test cases simulating implicit suicidal ideation in dialogue scenarios.
AUTALIC: A Dataset for Anti-AUTistic Ableist Language In Context (2025.acl-long)

Copied to clipboard

Challenge: Existing tools for detecting anti-autistic ableist language are lacking in this domain . current tools fail to accurately identify anti- autistic language, despite its subtle nature .
Approach: They present a dataset dedicated to the detection of anti-autistic ableist language in context . they use reddit sentences with surrounding context to identify anti-ableist expressions .
Outcome: AUTALIC is the first dataset dedicated to the detection of anti-autistic ableist language in context . it includes 2,400 autism-related sentences collected from reddit and annotated by trained experts .
Simulating Identity, Propagating Bias: Abstraction and Stereotypes in LLM-Generated Text (2025.findings-emnlp)

Copied to clipboard

Challenge: Persona-prompting is a growing strategy to personalize outputs, but its impact on how LLMs represent social groups remains underexplored.
Approach: They investigate whether persona-prompting leads to different levels of linguistic abstraction . they compare 11 persona driven responses to those of a generic AI assistant .
Outcome: The proposed method can be used to personalize outputs, but its impact on how LLMs represent social groups remains underexplored.
Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media (2025.acl-long)

Copied to clipboard

Challenge: Social media platforms are experiencing a growing presence of AI-Generated Texts (AIGTs) however, the misuse of AIGTs could have profound implications for public opinion .
Approach: They collect a dataset with 2.4M posts from 3 major social media platforms . they then construct a diverse dataset to train and evaluate AIGT detectors .
Outcome: The proposed dataset analyzes 2.4M posts from 3 major social media platforms from 2022 to 2024 . it finds that Medium and Quora show marked increases in AAR .
MASIVE: Open-Ended Affective State Identification in English and Spanish (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models that fail to understand cultural and language influences the meaning of emotional terms like "love" a new study shows that smaller finetuned models outperform much larger LLMs on region-specific span prediction tasks.
Approach: They propose to use a reddit reddits dataset to identify a set of affective states . they find that smaller finetuned multilingual models outperform larger LLMs .
Outcome: The proposed model outperforms larger models on span prediction task even on region-specific Spanish affective states.
Style-Shifting Behaviour of the Manosphere on Reddit (2024.emnlp-main)

Copied to clipboard

Challenge: Hate speech groups (HSGs) may negatively influence online platforms through their distinctive language, which may affect the tone and topic of discussion in other spaces if spread beyond the HSGs.
Approach: They explore the linguistic style of the Manosphere on reddit and how it reflects their linguistic styles across communities.
Outcome: The linguistic style of the Manosphere on Reddit is studied to determine whether it is harmful to health and community health.
STEntConv: Predicting Disagreement between Reddit Users with Stance Detection and a Signed Graph Convolutional Network (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to detect disagreements on social media platforms have focused on supplementing textual information with user network information, such as Twitter's following system, retweets and hashtags.
Approach: They propose a method which builds a graph of users and named entities and trains a Signed Graph Convolutional Network to detect disagreement between comment and reply posts.
Outcome: The proposed model builds a graph of users and named entities weighted by stance and trains a Signed Graph Convolutional Network (SGCN) to detect disagreement between comment and reply posts.
Dialogue Systems for Emotional Support via Value Reinforcement (2025.acl-long)

Copied to clipboard

Challenge: Emotional support dialogue systems aim to reduce help-seekers’ distress and help them overcome challenges.
Approach: They propose a value-driven method for training emotional support dialogue systems designed to reinforce positive values in seekers by leveraging online support conversations from Reddit.
Outcome: The proposed model outperforms baseline models across support skills, seekers’ emotional intensity, and value reinforcement.
A Multi-Level Benchmark for Causal Language Understanding in Social Media Discourse (2025.emnlp-main)

Copied to clipboard

Challenge: Existing datasets focus on explicit causality in structured text, providing limited support for detecting implicit causal expressions.
Approach: They propose a dataset of Reddit posts annotated across four causal tasks . they use a binary causal classification, explicit vs. implicit causality, cause–effect span extraction and causal gist generation to bridge causal detection and reasoning over informal discourse.
Outcome: The proposed dataset analyzes 10,120 Reddit posts discussing public health related to the COVID-19 pandemic.
PRIME: Large Language Model Personalization with Cognitive Dual-Memory and Personalized Thought Process (2025.emnlp-main)

Copied to clipboard

Challenge: Large language model (LLM) personalization aims to align outputs with individuals’ unique preferences and opinions.
Approach: They integrate a cognitive dual-memory model into LLM personalization by mirroring episodic memory to historical user engagements and semantic memory to long-term, evolving user beliefs.
Outcome: The proposed framework integrates the well-established cognitive dual-memory model into LLM personalization, using episodic and semanticmemories.
MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly being used by lay users for medical advice, but they have not yet been tested for this crucial competency.
Approach: They develop a semi-automated pipeline to curate MedRedFlag, a dataset of 1100+ reddit questions that require redirection.
Outcome: The proposed pipeline compares state-of-the-art LLMs to those from clinicians to find out how they perform under real-world health communication.
RubRIX: Rubric-Driven Risk Mitigation in Caregiver-AI Interactions (2026.findings-acl)

Copied to clipboard

Challenge: Existing frameworks focused on general risks may not adequately capture nuanced risks of LLMs in caregiving contexts.
Approach: They propose a theory-driven, clinician-validated framework for evaluating risks in LLMs . RubRIX operationalizes five empirically-derived risk dimensions: Inattention, Bias Stigma, Information Inaccuracy, Uncritical Affirmation, and Epistemic Arrogance.
Outcome: The proposed framework reduces risk components by 45-98% after one iteration across models.
FedMental: Evaluating Federated Learning for Mental Health Detection from Social Media Data (2026.acl-long)

Copied to clipboard

Challenge: Social media text data is often used to train machine learning models to identify users exhibiting high-risk mental health behaviors.
Approach: They apply federatedlearning and Differentially Private FL to two widely-studied mental health prediction tasks using social media text data.
Outcome: The proposed methods achieve comparable performance to centralized training on depression identification, but have a large performance-privacy trade-off even with low levels of noise.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations