Papers by David Jurgens
Copied to clipboard
| Challenge: | Empathy requires perspective-taking and is not explicitly modelled in NLP . |
| Approach: | They propose a new approach to recognizing alignment in empathetic speech, grounded in Appraisal Theory, and use reddit to study emotional conversations to examine alignment. |
| Outcome: | The proposed approach can recognize appraisals and alignments in empathetic speech, and mental health professionals engage with substantially more emotional alignment. |
Copied to clipboard
| Challenge: | Structured social variation has been extensively studied, e.g., gender based variation, but little is known about how to characterize individual styles due to their idiosyncratic nature. |
| Approach: | They propose a method to study idiolects through a massive cross-author comparison to identify and encode stylistic features. |
| Outcome: | The proposed model achieves strong performance at authorship identification on short texts and through an analogy-based probing task, showing that the learned representations exhibit surprising regularities that encode qualitative and quantitative shifts of idiolectal styles. |
Copied to clipboard
| Challenge: | Intimacy is a fundamental aspect of how we relate to others in social settings. |
| Approach: | They propose a computational framework for studying the intimacy in language with a dataset and a deep learning model for accurately predicting the intimacy level of questions. |
| Outcome: | The proposed framework enables the analysis of 80.5M questions across social media, books, and films to quantify the intimacy expressed in language and to predict the intimacy level. |
Copied to clipboard
| Challenge: | a dataset of over 1.1M podcast transcripts is largely comprehensive of all English language podcasts available through public RSS feeds from May and June of 2020. |
| Approach: | They propose to build a large-scale open dataset of podcast transcripts that includes metadata, speaker roles, audio features and speaker turns for a subset of 370K episodes. |
| Outcome: | The proposed dataset is largely comprehensive of all English language podcasts available through public RSS feeds from May and June of 2020. |
Copied to clipboard
| Challenge: | Prior work on linguistic gender difference and communications about gender has focused on language about or portraying persons of a particular gender. |
| Approach: | They present a multi-genre corpus of 25M comments from five socially and topically diverse sources tagged for the gender of the addressee and 30k annotations for sentiment and relevance of these responses. |
| Outcome: | The proposed dataset shows that responses to women are more emotive and about the speaker as an individual (rather than about the content being responded to). |
Copied to clipboard
| Challenge: | Citation context analysis (CCA) is an important task in natural language processing that studies how and why scholars discuss each other’s work. |
| Approach: | They propose to use a dataset of 12.6K citation contexts from 1.2K computational linguistics papers to model three important CCA phenomena. |
| Outcome: | The proposed dataset contains 12.6K citation contexts from 1.2K computational linguistics papers and can model these phenomena. |
Copied to clipboard
| Challenge: | lexical change is a prevalent process, as new words are added, thrive, and decline in day-to-day usage. |
| Approach: | They conduct a large-scale analysis of over 80k neologisms in 4420 online communities over a decade and found that the community’s network structure plays a significant role in lexical change. |
| Outcome: | The results show that the community’s network structure plays a significant role in lexical change. |
Copied to clipboard
| Challenge: | Existing tools for hate speech detection and sentiment analysis cannot detect veiled offensiveness of microaggressions . linguistic subtlety of micro-aggressives has made it difficult to analyze their exact nature . |
| Approach: | They propose a typology of microaggressions based on a subset of data . they propose an objective criterion for annotation and an active-learning procedure . |
| Outcome: | The proposed typology of microaggressions is based on a subset of social media data. |
Copied to clipboard
| Challenge: | Multilingual individuals code switch between languages as part of a complex communication process. |
| Approach: | They propose to model the social and contextual factors eliciting code switching in a rich contextual environment by analyzing 330K articles and 389K comments labeled for code switching behavior. |
| Outcome: | The proposed model shows that topic-driven variation, tribal affiliation, emotional valence, and audience design all play complementary roles in behavior. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are widely used to simulate human responses, but their ability to account for demographic differences in subjective tasks remains uncertain. |
| Approach: | They evaluate large language models' ability to understand demographic differences in two subjective judgment tasks: politeness and offensiveness. |
| Outcome: | The proposed models perform better in politeness and offensiveness tasks, while sociodemographic prompting does not improve and worsens their ability to perceive language from sub-populations. |
Copied to clipboard
| Challenge: | Whether the media faithfully communicate scientific information has long been a core issue to the science community. |
| Approach: | They propose to use the SCIENTIFIC PARAPHRASE AND INFORMATION CHANGE DATASET to identify paraphrased scientific findings annotated for degree of information change to enable large-scale tracking and analysis of information changes in science communication. |
| Outcome: | The proposed dataset contains 6,000 scientific finding pairs extracted from news stories, social media discussions, and full texts of original papers. |
Copied to clipboard
| Challenge: | Recent work treats disagreements as signal, instead of noise, resulting in a single label and marginalizing minoritized perspectives. |
| Approach: | They propose an approach to modeling annotator disagreement in subjective NLP tasks through architectural and data-centric innovations. |
| Outcome: | The proposed model performs competitively across demographic groups and shows strong results on datasets with high disagreement. |
Copied to clipboard
| Challenge: | Existing studies have found that presenting uncertainty in science communications influences people's perception of scientific findings and trust in science. |
| Approach: | They propose a model that models both the level and the aspects of certainty in scientific findings using an annotated dataset. |
| Outcome: | The proposed model can predict overall certainty and individual aspects of scientific findings with pre-trained language models, providing a more complete picture of the author’s intended communication. |
Copied to clipboard
| Challenge: | Academic citations are widely used for evaluating research and tracing knowledge flows. |
| Approach: | They propose a computational pipeline to quantify citation fidelity at the sentence level by identifying citations in citing papers and corresponding claims in cited papers. |
| Outcome: | The proposed pipeline identifies citations in citing papers and the corresponding claims in cited papers and applies supervised models to measure fidelity at the sentence level. |
Copied to clipboard
| Challenge: | Existing systems struggle to copy and properly cite unstructured evidence, which also tends to be “lost-in-the-middle”. |
| Approach: | They propose to extract unstructured evidence spans to improve the trustworthiness of large language models by citing unstructure . they propose to use this dataset as a training supervision for unstructure-based evidence summarization. |
| Outcome: | The proposed pipeline generates more relevant and factually consistent evidence than baselines with no fine-tuning and fixed granularity evidence. |
Copied to clipboard
| Challenge: | Personalized Large Language Models are increasingly used in diverse applications . prior research examined how well LLMs adhere to predefined personas in writing style . inconsistent responses are influenced by multiple factors, including the assigned persona, stereotypes, and model design choices. |
| Approach: | They propose a standardized framework to analyze consistency in persona-assigned LLMs. |
| Outcome: | The proposed framework evaluates personas across multiple tasks and runs. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are popular for research in social sciences . currently, prompting LLMs is insufficient to accurately and reliably capture model perceptions, and we discuss potential alternatives to improve this. |
| Approach: | They construct a dataset that contains 693 questions encompassing 39 different instruments of persona measurement on 115 persona axes and a set of questions containing minor variations. |
| Outcome: | The proposed model can generate answers and negate statements in a consistent and robust manner. |
Copied to clipboard
| Challenge: | Increasingly, image-based responses such as memes and animated gifs serve as culturally recognized and often humorous responses in conversation. |
| Approach: | They propose a multimodal conversational model for selecting gif responses from a text-gif conversation turn dataset and a randomized controlled trial. |
| Outcome: | The proposed model produces relevant and high-quality gif responses and is significantly better received by the community. |
Copied to clipboard
| Challenge: | In this paper, we analyze memes as a form of language subject to the same kinds of sociolinguistic variation as other modalities, such as written language and speech. |
| Approach: | They propose a computational pipeline to cluster memes into templates and semantic variables, taking advantage of their multimodal structure to learn meme semantics from an unstructured dataset. |
| Outcome: | The proposed method uses 3.8M images from a reddit meme database to analyze linguistic variation in memes. |
Copied to clipboard
| Challenge: | Empathy recognition and empathetic response generation tasks are well-established research directions, but there is little clarity on what empathy is and how it is being operationalized. |
| Approach: | They argue that current directions will benefit from a clear conceptualization that includes operationalizing cognitive empathy components. |
| Outcome: | The proposed framework will help to define and operationalize empathy in natural language processing. |
Copied to clipboard
| Challenge: | Occasionally, an event triggers a media storm, with coverage lasting weeks instead of days. |
| Approach: | They develop a pairwise article similarity model to identify story clusters in news corpora and build a corpus of media storms over a nearly two year period. |
| Outcome: | The proposed model validates theories about storm evolution and topical distribution and supports hypotheses about storm influence on media coverage and intermedia agenda setting. |
Copied to clipboard
| Challenge: | Authorship representation (AR) models capture an author's distinctive writing style by encoding documents written by the same author as nearby vectors in the embedding space. |
| Approach: | They propose a method that incorporates probabilistic content masking and language-aware batching to improve contrastive learning by reducing cross-lingual interference. |
| Outcome: | The proposed model outperforms monolingual baselines in 21 out of 22 non-English languages and reaches a maximum gain of 15.91% in a single language. |
Copied to clipboard
| Challenge: | Recent work has attempted to explain learning of style representations by generating natural language descriptions with large language models (LLMs) conditioned on input text. |
| Approach: | They propose a framework for interpreting style representations through style-eliciting prompts by prompting an LLM to generate text conditioned on these features. |
| Outcome: | The proposed framework outperforms baselines that directly prompt LLMs with target text, and achieves superior performance in both style description and style imitation. |
Copied to clipboard
| Challenge: | Psychological trauma can manifest following various distressing events, but studies focus on a single aspect of trauma, often neglecting the transferability of findings across different scenarios. |
| Approach: | They propose a language model that fine-tunes a single aspect of trauma to better predict traumatic events across domains. |
| Outcome: | The proposed model outperforms large language models on trauma-related datasets . it also outperformed models on court data, counseling conversations, and forum posts . |
Copied to clipboard
| Challenge: | Linguistic coordination is a phenomenon where conversation partners have similar patterns of language use. |
| Approach: | They propose a framework to organize the literature on linguistic coordination . they propose linguistic modeling choices and critiques of the choices involved . |
| Outcome: | The proposed framework provides an overview of the choices involved in the measurement process and synthesizes relevant critiques. |
Copied to clipboard
| Challenge: | Existing studies show that word choice is driven by demographics within the United States. |
| Approach: | They develop computational methods to study word choice within a sociolinguistic lexical variable . they use two variables to test for attitudes towards sexuality and gender in the u.s. |
| Outcome: | The proposed methods allow us to examine attitudes towards sexuality and gender in the United States through two lexical variables. |
Copied to clipboard
| Challenge: | POTATO is a free, fully open-sourced annotation system that supports labeling many types of text and multimodal data. |
| Approach: | They propose to use POTATO to design and deploy complex annotation tasks. |
| Outcome: | The proposed annotation system improves labeling speed and productivity over two tasks. |
Copied to clipboard
| Challenge: | Existing approaches to analyze moral reasoning are discordant and lack cohesion, focusing on isolated aspects of the process. |
| Approach: | They propose a unified dataset that integrates moral dilemmas annotated with labels for action choices, ethical principles, contributing factors, and consequences, and captures diverse socio-cultural contexts. |
| Outcome: | The proposed dataset integrates moral dilemmas annotated with labels for action choices, ethical principles, contributing factors, and consequences, along with annotators’ moral and cultural profiles. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly deployed in domains requiring moral understanding, yet their reasoning often remains shallow and misaligned with human reasoning. |
| Approach: | They propose a value-grounded framework for evaluating and distilling structured moral reasoning in large language models. |
| Outcome: | The proposed framework evaluates 12 open-source models across four moral datasets. |
Copied to clipboard
| Challenge: | Recent work has sought to use large language models to simulate human-human and human-LLM interactions. |
| Approach: | They use a large-scale dataset to generate a paired LLM-LLM and human-LLm dialogues from the WildChat dataset and quantify how well they align with their human counterparts. |
| Outcome: | The proposed models perform similarly in simulating English, Chinese, and Russian dialogues. |
Copied to clipboard
| Challenge: | Despite substantial efforts to reduce gender disparities in online social contexts, gender gaps persist and negatively affect women through online harassment. |
| Approach: | They propose a new dataset and method for identifying supportive replies and new methods for inferring gender from text and name to examine the disparity in support across millions of online interactions. |
| Outcome: | The proposed model shows that identifying as a woman is associated with higher rates of support, but also higher rates disparagement. |
Copied to clipboard
| Challenge: | Prior work has shown that partisan leanings can be inferred from a diverse set of behavioral characteristics such as text, social networks, and even community participation. |
| Approach: | They test this assumption and show that commonly-used models do not generalize . they also show that political users are more toxic on the platform and inter-party interactions are even more toxic . |
| Outcome: | The proposed models do not generalize, indicating heterogeneous political users. |
Copied to clipboard
| Challenge: | Variation in language is often linked to regional, social, and contextual factors. |
| Approach: | They propose a method to estimate tokenizer impact on downstream LLM performance . they pre-train BERT models with the popular Byte-Pair Encoding algorithm . |
| Outcome: | The proposed model improves on Rényi efficiency and other metrics on language variation. |
Copied to clipboard
| Challenge: | Using computational tools, we examine the dynamics of condolence online. |
| Approach: | They develop computational tools to analyze 11.4M distress expressions and 2.8M condolence offerings in a massive dataset of 11.4 million people. |
| Outcome: | The proposed model reveals that condolence features differ from those seen in interpersonal settings and that the features of condolance individuals find most helpful differ from the features seen in social media. |
Copied to clipboard
| Challenge: | Commercial AI systems often define the role of the LLM in system prompts. |
| Approach: | They conduct a systematic evaluation of personas in system prompts by adding 162 roles covering 6 types of interpersonal relationships and 8 domains of expertise. |
| Outcome: | The proposed model does not improve performance in the system prompt setting where no persona is added. |
Copied to clipboard
| Challenge: | Existing efforts to identify unacceptable behavior have focused on toxicity as the sole form of community norm violation. |
| Approach: | They propose a dataset that focuses on a more complete spectrum of community norms and their violations in local conversational and global contexts. |
| Outcome: | The proposed model improves the detection of community norm violations in local conversational and global contexts. |
Copied to clipboard
| Challenge: | Current methods to detect online abuse focus on a narrow definition of abuse to detriment of victims seeking validation and solutions. |
| Approach: | They argue that the NLP community needs to make three substantive changes to tackle both more subtle and more serious forms of abuse. |
| Outcome: | The proposed approach would address the problem of abuse in a more inclusive and productive way. |
Copied to clipboard
| Challenge: | Authorship verification (AV) is a crucial task for identity verification, accountlinking, historical linguistics, and AI-generated text identification. |
| Approach: | They propose to use Wikipedia's Million Authors Corpus to examine authorship verification models on a broad scale. |
| Outcome: | The proposed dataset includes 60.08M textual chunks, contributed by 1.29M Wikipedia authors. |
Copied to clipboard
| Challenge: | Tabular data is a foundational part of social sciences and is used to fit supervised learning models. |
| Approach: | They propose a technique for transforming tabular data to text data to improve deep learning models for tabular datasets. |
| Outcome: | The proposed technique improves performance of deep learning models for tabular data. |
Copied to clipboard
| Challenge: | a key intent behind many emails is to get a reply from the recipient. |
| Approach: | They propose to model the intents, expectations, and responsiveness in email exchanges by using a dataset containing 1800 emails annotated with nuanced types of intents and expectations. |
| Outcome: | The proposed model is based on 1800 emails annotated with nuanced types of intents and expectations . it shows that social status, argumentation, and strength of social connection influence email response rates . |
Copied to clipboard
| Challenge: | Annotated data is still central to NLP and Generative AI, yet the demands on annotation have grown in both scale and complexity. |
| Approach: | They introduce Potato 2.0, an open source annotation platform for easy deployment and customization. |
| Outcome: | The new version of potato supports 39 different types of annotation tasks and multiple AI-assistance features. |
Copied to clipboard
| Challenge: | Identifying educationally supportive contexts for vocabulary learning is an important problem to solve for designing effective curricula for contextual word learning. |
| Approach: | They evaluate attention-based approaches to find supportive contexts for vocabulary learning scenarios using an existing benchmark dataset. |
| Outcome: | The proposed model outperforms a generic model and a custom model on a major dataset for educational context support prediction. |
Copied to clipboard
| Challenge: | Existing benchmarks of social language are lacking for large language models. |
| Approach: | They propose a new benchmark that measures how well large language models understand social language by grouping 58 tasks into five categories: humor & sarcasm, offensiveness, sentiment & emotion, and trustworthiness. |
| Outcome: | The proposed model performs well at 58 tasks that are divided into five categories: humor & sarcasm, offensiveness, sentiment & emotion, and trustworthiness. |
Copied to clipboard
| Challenge: | Recent work suggests that annotators may have genuine disagreements, but few models separate signal from noise in annotator disagreement. |
| Approach: | They propose a Bayesian model that removes noisy annotations from training data while preserving systematic disagreements. |
| Outcome: | The proposed model outperforms models trained on NUTMEG-aggregated data. |
Copied to clipboard
| Challenge: | Recent work has shown that LLMs perform poorly when prompted with sociodemographic attributes, suggesting limited inherent sociodemography knowledge. |
| Approach: | They propose to train large language models to be accurate sociodemographic models of annotator variation by using a curated dataset of five tasks with standardized sociodemography. |
| Outcome: | The proposed models improve in sociodemographic prompting when trained but this performance gain is largely due to models learning annotator-specific behaviour rather than sociodemography. |
Copied to clipboard
| Challenge: | Empathy operationalizations in NLP are varied, with some having specific behaviors and properties, while others are more abstract. |
| Approach: | They analyze the transfer performance of empathy models adapted to empathy tasks with different theoretical groundings and characterize them as direct, abstract, or adjacent. |
| Outcome: | The proposed models show that they are more transferable than other models. |
Copied to clipboard
| Challenge: | VALUESCOPE is a framework that quantifies social norms and values within online communities. |
| Approach: | They propose a framework that uses language models to quantify social norms and values within online communities. |
| Outcome: | The proposed framework delineates differences in social norms and tracks evolution of norms in online communities and influence of significant external events like the U.S. presidential elections and the emergence of new sub-communities. |
Copied to clipboard
| Challenge: | Existing methods for identifying offensive content in interpersonal communication are largely independent of context . prior work has shown the benefits of modeling context, such as demographics of annotators and readers, and the online community in which a message is said. |
| Approach: | They propose a model that explicitly models the social context in which a message is said to assess whether it is appropriate. |
| Outcome: | The proposed model can accurately identify inappropriate communication in a given context. |
Copied to clipboard
| Challenge: | Using a dataset of immigration-related tweets, we examine how ordinary people on social media frame political issues. |
| Approach: | They propose to use a dataset of immigration-related tweets labeled for multiple framing typologies from political communication theory to analyze framers. |
| Outcome: | The proposed model enables comparisons between different types of frames on social media and a dataset of immigration-related tweets. |