Papers by Lyle Ungar
Building Knowledge-Guided Lexica to Model Cultural Variation (2024.naacl-long)
Copied to clipboard
| Challenge: | Cultural variation exists between nations, but also within regions . Historically, it has been difficult to computationally model cultural variation due to a lack of training data and scalability constraints. |
| Approach: | They propose a method to measure cultural variation using a knowledge-guided lexical model using geolocated tweets. |
| Outcome: | The proposed method could help us better understand the way people communicate and build more culturally-aware NLP systems. |
StyLEx: Explaining Style Using Human Lexical Annotations (2023.eacl-main)
Copied to clipboard
| Challenge: | Large pre-trained language models often learn spurious domain-specific words to make predictions. |
| Approach: | They propose a model that learns from human annotated explanations of stylistic features and jointly predicts them as model explanations. |
| Outcome: | The proposed model can provide human like stylistic lexical explanations without sacrificing performance on in-domain and out-of-domain datasets. |
An “Integrative Survey on Mental Health Conversational Agents to Bridge Computer Science and Medical Perspectives” (2023.emnlp-main)
Copied to clipboard
| Challenge: | Mental health conversational agents (a.k.a. chatbots) are widely studied for their potential to offer accessible support to those experiencing mental health challenges. |
| Approach: | They review 534 papers on building mental health-related conversational agents . they recommend a few recommendations to bridge the disciplinary divide . |
| Outcome: | The systematic review reveals 136 key papers on building mental health-related conversational agents with diverse characteristics of modeling and experimental design techniques. |
The Remarkable Benefit of User-Level Aggregation for Lexical-based Population-Level Predictions (D18-1)
Copied to clipboard
Salvatore Giorgi, Daniel Preoţiuc-Pietro, Anneke Buffone, Daniel Rieman, Lyle Ungar, H. Andrew Schwartz
| Challenge: | Social media data is often aggregated without regard to users in the Twitter populations of each community. |
| Approach: | They propose to use Twitter language to build community-level models using Twitter language aggregated by users. |
| Outcome: | The proposed method improves on four county-level tasks spanning demographic, health, and psychological outcomes over the standard approach of aggregating all tweets. |
Diachronic degradation of language models: Insights from social media (P18-2)
Copied to clipboard
| Challenge: | Existing studies have explored whether and how language models degrade over time, i.e. why they fail to work on contemporary language. |
| Approach: | They investigate the accuracy of pre-trained language models for downstream tasks in machine learning and user profiling. |
| Outcome: | The results show that it is possible to measure diachronic drifts within social media and within the span of a few years. |
Comparing Styles across Languages (2023.emnlp-main)
Copied to clipboard
| Challenge: | Communication practices vary across cultures. Inherent differences in how people think and behave influence cultural norms. |
| Approach: | They propose a framework to extract stylistic differences from multilingual language models (LMs) they use a multilingual lexica to consolidate feature importances into comparable lexical categories . |
| Outcome: | The proposed framework generates comprehensive style lexica in any language and consolidates feature importances from LMs into comparable lexical categories. |
Modeling Human Subjectivity in LLMs Using Explicit and Implicit Human Factors in Personas (2024.findings-emnlp)
Copied to clipboard
Salvatore Giorgi, Tingting Liu, Ankit Aich, Kelsey Isman, Garrick Sherman, Zachary Fried, João Sedoc, Lyle Ungar, Brenda Curtis
| Challenge: | Large language models (LLMs) are increasingly being used in human-centered social scientific tasks, such as data annotation, synthetic data creation, and engaging in dialog. |
| Approach: | They propose to prompt LLMs with human-like personas and ask them to answer as if they were a specific human, either explicitly, with exact demographics, political beliefs, and lived experiences, or implicitly via names prevalent in specific populations. |
| Outcome: | The proposed model is based on explicit, explicit, and implicit personas, and fails to show implicit biases. |
Modeling Empathy and Distress in Reaction to News Stories (D18-1)
Copied to clipboard
| Challenge: | a recent work on empathy prediction has underestimated the complexity of the phenomenon and lacks a shared corpus. authors present a novel annotation methodology which reliably captures empathy assessments by the writer of a statement using multi-item scales. |
| Approach: | They propose a method which captures empathy assessments by the writer of a statement using multi-item scales. |
| Outcome: | The proposed method distinguishes between multiple forms of empathy, empathic concern, and personal distress, as recognized throughout psychology. |
Towards Style Alignment in Cross-Cultural Translation (2025.acl-long)
Copied to clipboard
| Challenge: | Successful communication relies on the speaker’s intended style aligning with the listener’s interpreted style. |
| Approach: | They propose a method that leverages learned stylistic concepts to encourage LLM translation to appropriately convey cultural communication norms and align style. |
| Outcome: | The proposed method aims to encourage translations to convey cultural communication norms and align style. |
Social Norms in Cinema: A Cross-Cultural Analysis of Shame, Pride and Prejudice (2025.naacl-long)
Copied to clipboard
| Challenge: | We examine *how* and *why* shame and pride are expressed across cultures using a blend of psychology-informed language analysis combined with large language models. |
| Approach: | They introduce a cross-cultural dataset of over 10k shame/pride-related expressions with underlying social expectations from 5.4K Bollywood and Hollywood movies. |
| Outcome: | The results show that women are more sanctioned across cultures and for violating similar social expectations. |
A Holistic Framework for Analyzing the COVID-19 Vaccine Debate (2022.naacl-main)
Copied to clipboard
Maria Leonor Pacheco, Tunazzina Islam, Monal Mahajan, Andrey Shor, Ming Yin, Lyle Ungar, Dan Goldwasser
| Challenge: | Covid-19 infodemic has led to low quality information leading to poor health decisions . authors propose a framework for analyzing false claims and reasoning about the decisions a person makes . |
| Approach: | They propose a framework linking stance and reason analysis and moral sentiment analysis. |
| Outcome: | The proposed framework provides reliable predictions even in low-supervision settings. |
ControlText: Unlocking Controllable Fonts in Multilingual Text Rendering without Font Annotations (2025.findings-emnlp)
Copied to clipboard
Bowen Jiang, Yuan Yuan, Xinyi Bai, Zhuoqun Hao, Alyson Yin, Yaojie Hu, Wenyu Liao, Lyle Ungar, Camillo Jose Taylor
| Challenge: | a new method for visual text rendering requires glyph annotations to be obtained . |
| Approach: | They propose a model that integrates diffusion with a text segmentation model to achieve multilingual text rendering using just raw images without font label annotations. |
| Outcome: | The proposed model can achieve font-controllable multilingual text rendering without label annotations. |
Measuring the Language of Self-Disclosure across Corpora (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing models that estimate self-disclosure from language are poorly generalized due to variations in corpora and labeling instructions. |
| Approach: | They build single-task models on five self-disclosure corpora and use them to predict self-declaration across corpors. |
| Outcome: | The proposed model predicts self-disclosure across corpora, but the results are poor for out-of-corpora models. |
Characterizing Social Spambots by their Human Traits (2021.findings-acl)
Copied to clipboard
| Challenge: | Social spambots are an emerging class of spammers attempting to emulate people . previous studies show that standard spambot detection methods fail to distinguish them from genuine accounts . |
| Approach: | They hypothesize that human-like attributes of social spambots are unhuman-like . they find that social spam bots are extremely similar and average in their expressed personality, demographics, and emotion . |
| Outcome: | The proposed method is based on the human characteristics of social spambots . it shows that social bots are extremely similar and average in their expressed personality, demographics, and emotion . |
WikiTalkEdit: A Dataset for modeling Editors’ behaviors on Wikipedia (2021.naacl-main)
Copied to clipboard
| Challenge: | Using the WikiTalkEdit dataset, we show how positive emotion and the use of first-person pronouns predict a positive emotional change in a Wikipedia contributor. |
| Approach: | They introduce and analyze WikiTalkEdit, a dataset of conversations and edit histories from Wikipedia, for research in online cooperation and conversation modeling. |
| Outcome: | The proposed dataset supports the classic understanding of style matching, where positive emotion and the use of first-person pronouns predict a positive emotional change in a Wikipedia contributor. |
Conceptor-Aided Debiasing of Large Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Pre-trained large language models reflect inherent social biases of their training corpus. |
| Approach: | They propose two methods to identify and remove the bias subspace in pre-trained large language models such as BERT and GPT by applying conceptors to a conceptor NOT operation. |
| Outcome: | The proposed method achieves state-of-the-art (SoTA) debiasing while maintaining LLMs’ performance on the GLUE benchmark. |
ChatEval: A Tool for Chatbot Evaluation (N19-4)
Copied to clipboard
| Challenge: | open-domain dialog systems are difficult to evaluate due to lack of standardization and standardization in evaluation procedures. |
| Approach: | They propose a framework for human evaluation of chatbots that augments existing tools . researchers can submit their trained models to the ChatEval web interface . reproducibility and model assessment for opendomain dialog systems is challenging . |
| Outcome: | The proposed framework provides a web-based hub for researchers to compare their models with baselines and prior work. |
Does BERT Learn as Humans Perceive? Understanding Linguistic Styles through Lexica (2021.emnlp-main)
Copied to clipboard
| Challenge: | Using pre-trained models, people use different styles to express their interpersonal goal and attitude in their communication. |
| Approach: | They use a dataset to collect lexicon usages across styles using two lenses: human perception and machine word importance. |
| Outcome: | The proposed model can predict human perception and machine word importance based on a popular style classifier like BERT . human- and machine-identified words share significant overlap for some styles . |
Interactive Concept Learning for Uncovering Latent Themes in Large Text Collections (2023.findings-acl)
Copied to clipboard
| Challenge: | Topic modeling is a popular method for identifying emerging themes from text collections. |
| Approach: | They propose a framework that receives and encodes expert feedback at different levels of abstraction. |
| Outcome: | The proposed framework combines automation and manual coding, allowing experts to maintain control while reducing the manual effort required. |
Knowing When Not to Answer: Lightweight KB-Aligned OOD Detection for Safe RAG (2026.acl-long)
Copied to clipboard
| Challenge: | Retrieval-Augmented Generation (RAG) systems are increasingly deployed in high-stakes domains where safety depends on how a system answers . out-of-domain (OOD) queries can impair performance and safety . |
| Approach: | They propose to use lightweight, KB-aligned OOD detection as an always-on gate for RAG systems. |
| Outcome: | The proposed method scores queries in a compact subspace selected either by explained-variance retention (EVR) or by a separability-driven -test ranking. |
Predicting Responses to Psychological Questionnaires from Participants’ Social Media Posts and Question Text Embeddings (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing data cannot be used to predict responses for new questions or participants. |
| Approach: | They propose a method that uses social media texts and the text of the question to predict a participant's questionnaire response. |
| Outcome: | The proposed method can be used to integrate new participants or new questions into psychological studies without costly data collection. |
Language-based Valence and Arousal Expressions between the United States and China: a Cross-Cultural Examination (2025.findings-naacl)
Copied to clipboard
Young Min Cho, Dandan Pang, Stuti Thapa, Garrick Sherman, Lyle Ungar, Louis Tay, Sharath Chandra Guntuku
| Challenge: | valence and arousal are functionally equivalent across social media platforms . americans display higher emotional intensity than Chinese users . |
| Approach: | They compare valence and arousal on Twitter/X and Sina Weibo in China . they use the NRC-VAD lexicon to measure valance and valency . |
| Outcome: | The results show that the valence and arousal of the two platforms differ across cultures . the analysis also shows that the US users display higher emotional intensity than Chinese users . |
Systematic Evaluation of Auto-Encoding and Large Language Model Representations for Capturing Author States and Traits (2025.findings-acl)
Copied to clipboard
Khushboo Singh, Vasudha Varadarajan, Adithya V Ganesan, August Håkan Nilsson, Nikita Soni, Syeda Mahwish, Pranav Chitale, Ryan L. Boyd, Lyle Ungar, Richard N Rosenthal, H. Schwartz
| Challenge: | Large Language Models (LLMs) are increasingly used in human-centered applications, yet their ability to model diverse psychological constructs is not well understood. |
| Approach: | They evaluated a range of Transformer-LMs to predict psychological variables across five major dimensions: affect, substance use, mental health, sociodemographics, and personality. |
| Outcome: | The models predict affect, substance use, mental health, sociodemographics, and personality across five major dimensions. |
User-Level Race and Ethnicity Predictors from Twitter Text (C18-1)
Copied to clipboard
| Challenge: | Using social media text to identify user-level race and ethnicity is a useful tool for a range of downstream applications, including passive polling or quantifying demographic bias. |
| Approach: | They propose to collect data from social media users who self-report their race/ethnicity through a survey to develop models which accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC. |
| Outcome: | The proposed models accurately predict the membership of a user to the four largest racial and ethnic groups with up to .884 AUC and make available to the research community. |
The Impact of Language Mixing on Bilingual LLM Reasoning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies show multilingual speakers intentionally switch languages during reasoning . enforcing monolingual decoding reduces accuracy by 5.6 percentage points . |
| Approach: | They find that multilingual speakers intentionally switch languages during reasoning . enforcing monolingual decoding reduces accuracy by 5.6 percentage points . authors suggest that language mixing is not merely a byproduct of multilingual training . |
| Outcome: | The proposed model can be used to predict whether a language switch would benefit or harm reasoning. |
Inducing Generalizable and Interpretable Lexica (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Lexica are widely used as generalizable language features to predict sentiment, emotions, mental health, and personality. |
| Approach: | They propose to induce lexica using context-oblivious and context-aware approaches and compare their performance using crowd-worker assessment. |
| Outcome: | The proposed models can be induced using context-oblivious and context-aware approaches and evaluate their quality using crowd-worker assessment. |
Toward Micro-Dialect Identification in Diaglossic and Code-Switched Environments (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on dialect prediction is limited to coarse-grained varieties . a new language model, MARBERT, can predict micro-dialects with 9.9% F1, 76 better than a majority class baseline. |
| Approach: | They propose a new task of Micro-Dialect Identification (MDI) that can predict a fine-grained variety given a single message. |
| Outcome: | The proposed model predicts micro-dialects with 9.9% F1, 76 better than a majority class baseline. |
Identifying Locus of Control in Social Media Language (D18-1)
Copied to clipboard
| Challenge: | lexical features outperform syntactic features in expressing control in social media . authors communicate internal locus of control when they ascribe control to themselves . |
| Approach: | They examine the role of syntax and semantics in expressing users’ sense of control in annotated Facebook posts. |
| Outcome: | The proposed language outperforms syntactic features in identifying whether or not a user is in control of their circumstances. |
Conditioning on Dialog Acts improves Empathy Style Transfer (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing research has focused on empathetic response generation, but it is not applicable to more sensitive cases such as medicine and therapy where the content of the responses requires the supervision of medical experts. |
| Approach: | They propose two new style transfer strategies that use only examples of the target style and dialog-act-conditioned prompting to make a sentence more empathetic. |
| Outcome: | The proposed methods improve empathy more effectively while maintaining semantic similarity and preserving both semantics and the dialog-act type. |
Continual Learning for Sentence Representations Using Conceptors (N19-1)
Copied to clipboard
| Challenge: | Existing sentence encoders for distributed representations of sentences are limited in their performance on fixed corpora. |
| Approach: | They propose a continual learning scenario for distributed representations of sentences . they initialize sentence encoders with corpus-independent features and update them sequentially . |
| Outcome: | The proposed sentence encoder can learn features from new corpora while maintaining its competence on previously encountered corporales. |
Learning Word Ratings for Empathy and Distress from Document-Level User Responses (2020.lrec-1)
Copied to clipboard
| Challenge: | Emotion analysis of text is increasing in popularity in NLP, however, manually creating lexica for psychological constructs such as empathy has proven difficult. |
| Approach: | They compare different approaches to learning word ratings from higher-level supervision and use a Mixed-Level Feed Forward Network to create the first-ever empathy lexicon. |
| Outcome: | The proposed model automatically creates empathy word ratings from document-level ratings. |
Unsupervised Morphology Learning with Statistical Paradigms (C18-1)
Copied to clipboard
| Challenge: | Existing models treat words as concatenation of morphemes, but some use transformations like rewrite rules to recognize dependencies between morphs. |
| Approach: | They propose an unsupervised model that exploits the notion of paradigms for morphological segmentation that can be applied to a homogeneous set of words. |
| Outcome: | The proposed model significantly improves on the Morpho-Challenge dataset in English, Turkish, and Finnish. |