Papers by Rob Voigt
Good Intentions Beyond ACL: Who Does NLP for Social Good, and Where? (2025.emnlp-main)
Copied to clipboard
| Challenge: | 20% of all papers in the ACL Anthology address social good issues . authors are more likely to do work addressing social good concerns when publishing in venues outside of ACL. |
| Approach: | They use author- and venue-level perspectives to map the landscape of NLP4SG . they find authors are more likely to do work addressing social good concerns outside of ACL . |
| Outcome: | The study analyzes the literature on NLP4SG and its impact on the ACL community . 20% of all papers in the anthology address social good issues, the study finds . |
RtGender: A Corpus for Studying Differential Responses to Gender (L18-1)
Copied to clipboard
| Challenge: | Prior work on linguistic gender difference and communications about gender has focused on language about or portraying persons of a particular gender. |
| Approach: | They present a multi-genre corpus of 25M comments from five socially and topically diverse sources tagged for the gender of the addressee and 30k annotations for sentiment and relevance of these responses. |
| Outcome: | The proposed dataset shows that responses to women are more emotive and about the speaker as an individual (rather than about the content being responded to). |
Thinking Out Loud: Do Reasoning Models Know When They’re Right? (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large reasoning models (LRMs) have recently demonstrated impressive capabilities in complex reasoning tasks by leveraging increased test-time computation and exhibiting behaviors reminiscent of human-like self-reflection. |
| Approach: | They analyze verbalized confidence, how models articulate their certainty, as a lens into the nature of self-reflection in large reasoning models. |
| Outcome: | The proposed model exhibits human-like self-reflection in reasoning tasks, but how this ability interacts with other model behaviors remains underexplored. |
A Computational Analysis and Exploration of Linguistic Borrowings in French Rap Lyrics (2024.acl-srw)
Copied to clipboard
| Challenge: | rap is a popular genre in the u.s. and has been used in countries far beyond the uk . linguistic borrowings are especially intriguing in countries such as the eu and europe . |
| Approach: | They manually annotate a lexicon of over 700 borrowings in the French language . they find that there are increases in the proportion of linguistic borrowings, interjections, and Niger-Congo borrowings . |
| Outcome: | The proposed method analyzes a corpus of over 8000 french rap song lyrics and shows that rap borrowings are increasing in prevalence and interjections are decreasing. |
Socially Responsible NLP (N18-6)
Copied to clipboard
| Challenge: | This tutorial will provide an overview of ethical research tools and ethical implications of language technologies. |
| Approach: | This tutorial will provide an overview of ethical research and practical examples . it will discuss ethical tools to ensure data, algorithms, and models are socially responsible . |
| Outcome: | This tutorial will provide an overview of ethical research tools and methods . it will discuss philosophical foundations of ethical work along with state of the art techniques . |
Language of Bargaining (2023.acl-long)
Copied to clipboard
| Challenge: | a new dataset is being developed to study how language shapes bilateral bargaining . a recent study examined the use of language in negotiation education . |
| Approach: | They propose a dataset to study how language shapes bilateral bargaining . they recruit participants via behavioral labs instead of crowdsourcing platforms . |
| Outcome: | The proposed dataset is based on an exercise in negotiation education . it shows that when subjects can talk, negotiations finish faster and prices drop . |
The Pragmatic Mind of Machines: Tracing the Emergence of Pragmatic Competence in Large Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Current large language models (LLMs) have demonstrated emerging capabilities in social intelligence tasks, including implicature resolution and theory-of-mind reasoning. |
| Approach: | They introduce a dataset grounded in the pragmatic concept of alternatives to evaluate whether large language models can accurately infer nuanced speaker intentions. |
| Outcome: | The proposed model can infer nuanced speaker intentions by inferring the speaker’s intended meaning and explaining when and why a speaker would choose one utterance over its alternative. |
Adaptive Axes: A Pipeline for In-domain Social Stereotype Analysis (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to quantify social stereotypes have struggled to capture the variability in stereotypes across conceptual domains for the same social group. |
| Approach: | They propose to use text embedding models and adaptive semantic axes to recover stereotypes from contextual representations by using large language models. |
| Outcome: | The proposed pipeline surpasses token-based methods in capturing in-domain framing and tracks stereotypes along domain-specific semantic axes for in- domain texts. |
Leveraging Human Production-Interpretation Asymmetries to Test LLM Cognitive Plausibility (2025.acl-short)
Copied to clipboard
| Challenge: | Existing research on the linguistic capabilities of large language models has focused on their performance in language interpretation. |
| Approach: | They examine whether large language models (LLMs) process language similarly to humans . they use an empirically documented asymmetry between production and interpretation in humans a testbed . |
| Outcome: | The proposed model can replicate human-like distinctions between production and interpretation. |
Large Language Models Are Partially Primed in Pronoun Interpretation (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing studies suggest large language models acquire rich linguistic representations, but little is known about whether they adapt to linguistic biases in a human-like way. |
| Approach: | They examine whether large language models display human-like referential biases using stimuli and procedures from real psycholinguistic experiments. |
| Outcome: | The proposed models display human-like referential biases when exposed to referential patterns in the local context. |
Analyzing Polarization in Social Media: Method and Application to Tweets on 21 Mass Shootings (N19-1)
Copied to clipboard
| Challenge: | a new framework for studying political polarization in social media is needed to understand how group divisions manifest in language. |
| Approach: | They propose to cluster tweet embeddings to uncover four dimensions of political polarization in social media . their results apply existing lexical methods to analyze 4.4M tweets on 21 mass shootings . |
| Outcome: | The proposed framework generates more cohesive topics than traditional models. |