Papers by Caleb Ziems
Culture Cartography: Mapping the Landscape of Cultural Knowledge (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can empower users to be more knowledgeable, productive, and creative, but their utility is often diminished for under-represented groups and cultures. |
| Approach: | They propose a methodology that operationalizes a mixed-initiative approach to finding culture-specific knowledge that is salient to in-group users but unknown to LLMs. |
| Outcome: | The proposed method improves the accuracy of LLMs on culturally-competent language models by 19.2%. |
To Protect and To Serve? Analyzing Entity-Centric Framing of Police Violence (2021.findings-emnlp)
Copied to clipboard
| Challenge: | a new study examines the media coverage of police violence in the United States by examining the framing of 82k news articles spanning 7k police killings. |
| Approach: | They propose an NLP framework to measure entity-centric framing to understand media coverage on police violence in the United States in a new police violence frames corpus of 82k news articles spanning 7k police killings. |
| Outcome: | The proposed framework reveals significant differences in the way liberal and conservative news sources frame both the issue of police violence and the entities involved. |
Modeling Cross-Cultural Pragmatic Inference with Codenames Duet (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing work on pragmatic reasoning tests using simple word reference games with unidentified speakers and listeners, but speakers' sociocultural background shapes their pragmatic assumptions. |
| Approach: | They propose a dataset which operationalizes sociocultural pragmatic inference in a word reference game. |
| Outcome: | The proposed model improves clue-giving and guessing tasks by accounting for background characteristics and the game context. |
NormBank: A Knowledge Bank of Situational Social Norms (2023.acl-long)
Copied to clipboard
| Challenge: | NormBank is a knowledge bank of 155k situational norms that can be used to ground flexible normative reasoning for interactive, assistive, and collaborative AI systems. |
| Approach: | They propose a new scheme for hierarchically organizing the seemingly unbounded social norms within a multivalent sociocultural frame. |
| Outcome: | The proposed framework can be used to ground flexible reasoning for interactive, assistive, and collaborative AI systems. |
The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems (2022.acl-long)
Copied to clipboard
| Challenge: | Moral integrity corpus captures the moral assumptions of 38k prompt-reply pairs, using 99k distinct Rules of Thumb (RoTs). |
| Approach: | They propose a resource that captures the moral assumptions of 38k prompt-reply pairs, using 99k distinct Rules of Thumb (RoTs). |
| Outcome: | The proposed resource captures the moral assumptions of 38k prompt-reply pairs, using 99k distinct Rules of Thumb (RoTs). |
Multi-VALUE: A Framework for Cross-Dialectal English NLP (2023.acl-long)
Copied to clipboard
| Challenge: | Current systems that focus on standard American English are not dialect invariant . current systems focus on a single dialect, which results in performance discrepancies . |
| Approach: | They propose a resource for evaluating and achieving English dialect invariance . they stress test question answering, machine translation, and semantic parsing . |
| Outcome: | The proposed system is based on a rule-based translation system spanning 50 English dialects and 189 unique linguistic features. |
TADA : Task Agnostic Dialect Adapters for English (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing work on dialectal English NLP is task-specific, using manual annotated dialect data, weak supervision, or data augmentation. |
| Approach: | They propose a method for task-agnostic dialect adaptation by aligning non-SAE dialects with task-specific adapters from SAE. |
| Outcome: | The proposed method improves dialectal robustness on 4 dialectal variants of the GLUE benchmark without task-specific supervision. |
VALUE: Understanding Dialect Disparity in NLU (2022.acl-long)
Copied to clipboard
| Challenge: | English Natural Language Understanding systems outperform humans on benchmarks like GLUE and SuperGLUE, but they only use textbook Standard American English (SAE) . fewer studies have considered the effects of dialectal differences on performance . |
| Approach: | They propose a benchmark to evaluate the performance of English natural language understanding systems using a set of lexical and morphosyntactic transformation rules. |
| Outcome: | The proposed model outperforms humans on GLUE and SuperGLUE, but only on standard American English . the proposed model recruits fluent speakers of African American vernacular english to validate each feature transformation . |
Social Intelligence Data Infrastructure: Structuring the Present and Navigating the Future (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing work on social intelligence in NLP does not provide a coherent subfield for researchers to analyze and identify research gaps and future directions. |
| Approach: | They build a social AI taxonomy and a data library of 480 NLP datasets to analyze existing datasets and evaluate language models’ performance in different social intelligence aspects. |
| Outcome: | The proposed infrastructure analyzes existing dataset efforts and evaluates language models’ performance in different social intelligence aspects. |
Measuring and Addressing Indexical Bias in Information Retrieval (2024.findings-acl)
Copied to clipboard
| Challenge: | Information Retrieval (IR) systems may not optimize rankings for fairness, neutrality, or the balance of ideas. |
| Approach: | They propose to use a framework to automatically audit IR rankings for indexical biases, or biase in the positional order of documents. |
| Outcome: | The proposed bias metric can help predict when and how indexical bias will shift a reader’s opinion. |
Silent Signals, Loud Impact: LLMs for Word-Sense Disambiguation of Coded Dog Whistles (2024.acl-long)
Copied to clipboard
| Challenge: | a dog whistle is a coded communication that carries a secondary meaning to specific audiences and is often weaponized for racial and socioeconomic discrimination. |
| Approach: | They propose an approach for word-sense disambiguation of dog whistles from standard speech using Large Language Models. |
| Outcome: | The proposed method allows disambiguation of dog whistles from standard speech using large language models. |
Impressions: Visual Semiotics and Aesthetic Impact Understanding (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing image captioning and conditional generation models struggle to simulate plausible human responses to images. |
| Approach: | They propose a dataset to investigate the semiotics of images and how visual features and design choices can elicit specific emotions, thoughts and beliefs. |
| Outcome: | The proposed dataset improves existing models for image captioning and conditional generation. |
Inducing Positive Perspectives with Text Reframing (2022.acl-long)
Copied to clipboard
| Challenge: | Sentiment transfer is a text style transfer task that aims to reverse sentiment polarity and reversal in meaning. |
| Approach: | They propose a task called positive reframing that neutralizes a negative point of view and generates 'positive' perspectives without contradicting original meaning. |
| Outcome: | The proposed model neutralizes a negative point of view and generates 'positive' perspectives without contradicting the original meaning. |
Latent Hatred: A Benchmark for Understanding Implicit Hate Speech (2021.emnlp-main)
Copied to clipboard
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, Diyi Yang
| Challenge: | Existing studies on explicit or overt hate speech have failed to address a more pervasive form based on coded or indirect language. |
| Approach: | They propose a theoretically-justified taxonomy of implicit hate speech and a benchmark corpus with fine-grained labels for each message and its implication. |
| Outcome: | The proposed dataset will serve as a useful benchmark for understanding this multifaceted issue. |
CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data Annotation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Annotated data plays a critical role in training models and evaluating their performance. |
| Approach: | They propose a paradigm for Human-LLM co-annotation of unstructured texts at scale that utilizes uncertainty to estimate LLMs’ annotation capability. |
| Outcome: | The proposed model outperforms existing models on many text-annotation tasks with up to 21% performance improvement over random baseline. |
CultureBank: An Online Community-Driven Knowledge Base Towards Culturally Aware Language Technologies (2024.findings-emnlp)
Copied to clipboard
| Challenge: | CultureBank is a knowledge base built upon users’ self-narratives with 12K cultural descriptors sourced from TikTok and 11K from Reddit. |
| Approach: | They construct a pipeline to construct cultural knowledge bases from different online communities on a massive scale. |
| Outcome: | The proposed pipeline improves cultural awareness of language models by evaluating them on two cultural tasks in a zero-shot setting. |
EgoNormia: Benchmarking Physical-Social Norm Understanding (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing VLMs lack robust grounded norm understanding, a new study finds . current VLM models lack robust grounding, despite a high score for safety and privacy . |
| Approach: | They propose a pipeline to generate grounded MCQs from ego-centric videos of human interactions. |
| Outcome: | The proposed pipeline can generate grounded MCQs from egocentric video . it shows that current VLMs lack robust grounded norm understanding . |