Papers by Vered Shwartz
Uncovering Implicit Gender Bias in Narratives through Commonsense Inference (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Pre-trained language models learn harmful biases from their training corpora and may repeat these biase if used for generation. |
| Approach: | They focus on gender biases associated with the protagonist in model-generated stories and use a commonsense reasoning engine to uncover them. |
| Outcome: | The proposed model-generated stories are based on a commonsense reasoning engine and are able to uncover gender biases in the protagonist's motivations, attributes, mental states, and implications on others. |
Thinking Like a Skeptic: Defeasible Inference in Natural Language (2020.findings-emnlp)
Copied to clipboard
Rachel Rudinger, Vered Shwartz, Jena D. Hwang, Chandra Bhagavatula, Maxwell Forbes, Ronan Le Bras, Noah A. Smith, Yejin Choi
| Challenge: | Defeasible inference is a mode of reasoning in which an inference may be weakened or overturned in light of new evidence. |
| Approach: | They propose a dataset for defeasible inference in natural language that includes extensions to existing inference datasets. |
| Outcome: | Defeasible NLI extends existing datasets for defeaasibility inference in natural language . generative models can weaken or strengthen inferences up to 68% of the time, it shows . |
Revisiting Joint Modeling of Cross-document Entity and Event Coreference Resolution (P19-1)
Copied to clipboard
| Challenge: | Recognizing that various textual spans across multiple texts refer to the same entity or event is an important NLP task. |
| Approach: | They propose a neural architecture for cross-document coreference resolution by representing an event mention using its lexical span, surrounding context, and relation to other mentions via predicate-arguments structures. |
| Outcome: | The proposed model outperforms the state-of-the-art event coreference model on ECB+ while providing the first entity coreference results on this corpus. |
Evaluating Text GANs as Language Models (N19-1)
Copied to clipboard
| Challenge: | Generative Adversarial Networks (GANs) do not suffer from the problem of exposure bias. |
| Approach: | They propose to approximate the distribution of text generated by a GAN and compare it to traditional probability-based LM metrics. |
| Outcome: | The proposed method performs significantly worse than state-of-the-art LMs on several GAN-based models and can accelerate progress in GAN text generation. |
Unsupervised Commonsense Question Answering with Self-Talk (2020.emnlp-main)
Copied to clipboard
| Challenge: | Current systems rely on pre-trained language models or external knowledge bases to incorporate additional relevant knowledge. |
| Approach: | They propose an unsupervised framework based on self-talk to improve commonsense performance by asking language models to ask information seeking questions. |
| Outcome: | Empirical results show that the proposed framework improves on four out of six commonsense benchmarks and competes with models that obtain knowledge from external KBs. |
It’s not Rocket Science: Interpreting Figurative Language in Narratives (2022.tacl-1)
Copied to clipboard
| Challenge: | Existing text representations by design rely on compositionality, while figurative language is often non-compositional. |
| Approach: | They propose to use a pre-trained language model to interpret figurative language types to adopt human strategies for interpreting figurativ language types: inferring meaning from context and relying on constituent words’ literal meanings. |
| Outcome: | The proposed models perform significantly worse than humans on discriminative and generative tasks, bridging the gap from human performance. |
From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Vision-Language Models (VLMs) have shown emerging capabilities through large-scale training that have made them gain popularity in recent years. |
| Approach: | They propose to perform retrieval across universals and cultural visual grounding tasks to assess cultural diversity across universal and culture-specific local concepts. |
| Outcome: | The proposed benchmarks show that the models perform significantly across cultures, underscoring the need for enhancing multicultural understanding in vision-language models. |
From chocolate bunny to chocolate crocodile: Do Language Models Understand Noun Compounds? (2023.findings-acl)
Copied to clipboard
| Challenge: | Noun compound interpretation is the task of expressing a noun compound in a free-text paraphrase that makes the relationship between the constituent nouns explicit. |
| Approach: | They propose modifications to the standard task and propose a new task that solves it. |
| Outcome: | The proposed task solves the standard task of paraphrasing a noun compound in a free-text paraphrase that makes the relationship between the constituent nouns explicit. |
Small But Funny: A Feedback-Driven Approach to Humor Distillation (2024.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been used to transfer knowledge from LLMs to smaller, smaller language models (SLMs). |
| Approach: | They propose to assign a dual role to the LLM as a “teacher” generating data, as well as evaluating the student’s performance. |
| Outcome: | The proposed approach narrows the performance gap between LLMs and larger models by incorporating feedback into the data. |
MemeCap: A Dataset for Captioning and Interpreting Memes (2023.emnlp-main)
Copied to clipboard
| Challenge: | a new dataset aims to understand meme captioning tasks using visual metaphors . vision and language models are proving to be effective in image captioning and visual question answering tasks . |
| Approach: | They present a dataset that contains 6.3K memes and 6.3k meme captions . they show that vision and language models still struggle with visual metaphors despite their advanced capabilities . |
| Outcome: | The proposed dataset contains 6.3K memes along with the title of the post containing the meme, meme captions, literal image caption, and visual metaphors. |
CulturalBench: A Robust, Diverse and Challenging Benchmark for Measuring LMs’ Cultural Knowledge Through Human-AI Red-Teaming (2025.acl-long)
Copied to clipboard
Yu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, Yejin Choi
| Challenge: | CulturalBench is a set of 1,696 human-written and human-verified questions to assess LMs’ cultural knowledge covering 45 global regions including underrepresented ones like Bangladesh, Zimbabwe, and Peru. |
| Approach: | They construct a set of 1,696 human-written and human-verified questions to assess LMs' cultural knowledge, covering 45 global regions including underrepresented ones like Bangladesh, Zimbabwe, and Peru. |
| Outcome: | The proposed model outperforms other models across cultures, while underperforming on questions related to North Africa, South America and Middle East. |
Good Night at 4 pm?! Time Expressions in Different Cultures (2022.findings-acl)
Copied to clipboard
| Challenge: | a new paper focuses on temporal grounding of time expressions to specific hours in the day . we propose language-agnostic methods for mapping time expression to specific times . cultural differences can cause variation in interpretation of time-specific expressions . |
| Approach: | They propose to use language-agnostic methods to map time expressions to specific hours . they use a time expression that is interpreted by different people . |
| Outcome: | The proposed method achieves promising results on gold standard annotations for 27 languages. |
Paraphrase to Explicate: Revealing Implicit Noun-Compound Relations (P18-1)
Copied to clipboard
| Challenge: | Existing methods for paraphrasing nouncompounds lack the ability to generalize and have a hard time interpreting infrequent or new noun-compound. |
| Approach: | They propose a neural model that generalizes better by representing paraphrases in a continuous space, generalizing for both unseen noun-compounds and rare paraphrase. |
| Outcome: | The proposed model generalizes better by representing paraphrases in a continuous space, generalizing for unseen noun-compounds and rare paraphrase. |
Social Chemistry 101: Learning to Reason about Social and Moral Norms (2020.emnlp-main)
Copied to clipboard
| Challenge: | SOCIAL CHEMISTRY is a conceptual formalism to study people’s everyday social norms and moral judgments over a rich spectrum of real life situations described in natural language. |
| Approach: | They propose a new conceptual formalism to study people's everyday social norms and moral judgments over a rich spectrum of real life situations described in natural language. |
| Outcome: | The proposed model can be used to model people's everyday social norms and moral judgments over a rich spectrum of real life situations. |
Teach the Rules, Provide the Facts: Targeted Relational-knowledge Enhancement for Textual Inference (2021.starsem-1)
Copied to clipboard
| Challenge: | InferBERT is a method to enhance transformer-based inference models with relevant relational knowledge. |
| Approach: | They propose to enhance transformer-based inference models with relevant relational knowledge by injecting relevant facts at test time into the model. |
| Outcome: | The proposed method outperforms existing models on the challenge datasets while outperforming existing models. |
Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models (2024.eacl-long)
Copied to clipboard
Natalie Shapira, Mosh Levy, Seyed Hossein Alavi, Xuhui Zhou, Yejin Choi, Yoav Goldberg, Maarten Sap, Vered Shwartz
| Challenge: | Recent work suggests that Large Language Models (LLMs) exhibit Neural Theory-of-Mind (N-ToM) however, prior work reached conflicting conclusions regarding those abilities. |
| Approach: | They examine the extent of Large Language Models’ N-ToM abilities through an extensive evaluation of 6 tasks and find that LLMs struggle with adversarial examples . |
| Outcome: | The proposed metrics show that LLMs exhibit certain N-ToM abilities, but this behavior is far from robust. |
Olive Oil is Made of Olives, Baby Oil is Made for Babies: Interpreting Noun Compounds Using Paraphrases in a Neural Model (N18-2)
Copied to clipboard
| Challenge: | Recent work suggests that success stems from memorizing single prototypical words for each relation. |
| Approach: | They propose a neural paraphrasing approach that maps NCs to paraphrases that express the relation between constituent words. |
| Outcome: | The proposed method performs better when memorization is not possible. |
A Graph per Persona: Reasoning about Subjective Natural Language Descriptions (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing large language models (LLMs) perform poorly in reasoning about subjective knowledge, showing strong biases and lack interpretability requirements. |
| Approach: | They propose a novel approach for reasoning about subjective knowledge that integrates potential and implicit meanings and explicitly models the relational nature of the information. |
| Outcome: | The proposed model outperforms several prominent large language models on the OpinionQA dataset, showing its unique advantages and complementary nature. |
Breaking NLI Systems with Sentences that Require Simple Lexical Inferences (P18-2)
Copied to clipboard
| Challenge: | a new test set shows the deficiency of state-of-the-art models in inferences that require lexical and world knowledge. |
| Approach: | They create a new NLI test set that shows the deficiency of state-of-the-art models in inferences that require lexical and world knowledge. |
| Outcome: | The new examples are simpler than the SNLI test set, but the state-of-the-art systems perform poorly on it. |
Commonsense Reasoning for Natural Language Processing (2020.acl-tutorials)
Copied to clipboard
| Challenge: | In this tutorial, we will outline the various types of commonsense knowledge and discuss techniques to gather and represent commonsence knowledge. |
| Approach: | This tutorial will provide researchers with the critical foundations and recent advances in commonsense representation and reasoning. |
| Outcome: | This tutorial will outline the various types of commonsense and discuss techniques to gather and represent commonsence knowledge while highlighting the challenges specific to this type of knowledge (e.g., reporting bias). |
Locating Information Gaps and Narrative Inconsistencies Across Languages: A Case Study of LGBT People Portrayals on Wikipedia (2024.emnlp-main)
Copied to clipboard
| Challenge: | a recent study focuses on comparative text analyses to explain social phenomena and identify systematic biases. |
| Approach: | They evaluate InfoGap method to locate information gaps and inconsistencies in articles at the fact level, across languages. |
| Outcome: | The method identifies discrepancies in factual coverage across languages and biographical facts carrying negative connotations are more likely to be highlighted in Russian Wikipedia. |
Do Neural Language Models Overcome Reporting Bias? (2020.coling-main)
Copied to clipboard
| Challenge: | Recent studies show that pre-trained language models can overcome reporting bias by estimating the plausibility of rare but unspoken facts. |
| Approach: | They revisit the experiments conducted by Gordon and Van Durme (2013) . they find that pre-trained language models overestimate the very rare . |
| Outcome: | The proposed approach overestimates the rare at the expense of the rare, while minimizing reporting bias. |
CASE: Commonsense-Augmented Score with an Expanded Answer Space (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated impressive zero-shot performance on NLP tasks thanks to the knowledge they acquired in their training. |
| Approach: | They propose a Commonsense-Augmented Score with an Expanded Answer Space that assigns importance weights to words based on their semantic relations to other words in the input. |
| Outcome: | The proposed approach outperforms basic LM scores on 5 commonsense benchmarks and is complementary to previous approaches. |
COMET-M: Reasoning about Multiple Events in Complex Sentences (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing commonsense models that generate event-centric inferences for simple sentences struggle with the complexity of multi-event sentences prevalent in natural text. |
| Approach: | They propose a commonsense model that generates inferences for a target event within a complex sentence using a multi-event inference dataset. |
| Outcome: | The proposed model produces inferences for a target event within a complex sentence taking the complete context into account. |
Back to the Future: Unsupervised Backprop-based Decoding for Counterfactual and Abductive Commonsense Reasoning (2020.emnlp-main)
Copied to clipboard
Lianhui Qin, Vered Shwartz, Peter West, Chandra Bhagavatula, Jena D. Hwang, Ronan Le Bras, Antoine Bosselut, Yejin Choi
| Challenge: | Existing methods for integrating past and future contexts are limited and require manual input. |
| Approach: | They propose an unsupervised decoding algorithm that incorporates past and future contexts using off-the-shelf, left-to-right language models and no supervision. |
| Outcome: | The proposed method outperforms unsupervised methods on abductive and counterfactual reasoning tasks. |
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities (2025.acl-long)
Copied to clipboard
| Challenge: | Vision-language Models have been shown to be highly capable but lacking basic visual understanding skills. |
| Approach: | They propose to examine the limitations of vision-language models on visual tasks by constructing a series of tests that probe which components of design may be lacking. |
| Outcome: | The proposed tests compare VLMs to other models on visual encoders, intermediate vision-language projection and LLM-decoder outputs. |
“You are grounded!”: Latent Name Artifacts in Pre-trained Language Models (2020.emnlp-main)
Copied to clipboard
| Challenge: | Pre-trained language models perpetuate biases originating in their training corpus to downstream models. |
| Approach: | They focus on the representations of given names in pre-trained language models and show that name perturbation can have an effect on downstream tasks. |
| Outcome: | The proposed model can be used to model the representation of given names in pre-trained language models on reading comprehension probes where name perturbation changes the model answers. |
Infusing Theory of Mind into Socially Intelligent LLM Agents (2026.findings-acl)
Copied to clipboard
| Challenge: | Theory of Mind (ToM) is a key aspect of human social intelligence, yet chatbots and LLMs do not typically integrate it. |
| Approach: | They propose a method that integrates Theory of Mind (ToM) into chatbots and dialogue agents to generate mental states between dialogue turns. |
| Outcome: | The proposed method improves dialogue and social interaction by integrating ToM with dialogue lookahead. |
GD-COMET: A Geo-Diverse Commonsense Inference Model (2023.emnlp-main)
Copied to clipboard
| Challenge: | GD-COMET is a geo-diverse version of the COMET commonsense inference model . it captures and generates culturally nuanced commonsensense knowledge . lack of cultural awareness may lead to models perpetuating stereotypes and reinforcing societal inequalities for users from non-Western countries. |
| Approach: | They propose a geo-diverse version of COMET commonsense reasoning model that generates inferences pertaining to a broad range of cultures. |
| Outcome: | The proposed model generates inferences pertaining to a broad range of cultures and is culturally nuanced. |
What happens before and after: Multi-Event Commonsense in Event Coreference Resolution (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing event coreference models cluster event mentions pertaining to the same event, but they fail to leverage commonsense inferences for lexically-divergent mentions. |
| Approach: | They propose a model that extends event mentions with temporal commonsense inferences to generate plausible events that happen before and after the target events. |
| Outcome: | The proposed model generates plausible events that happen before and after the target event, and then after it, such as "he was sentenced". |
Knowledge Graph Compression Enhances Diverse Commonsense Generation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing models use commonsense knowledge graphs to extract subgraphs of relevant knowledge pertaining to concepts in the input but due to the large coverage and vast scale of ConceptNet, the extracted subgraph may contain loosely related, redundant and irrelevant information. |
| Approach: | They propose to apply a differentiable graph compression algorithm to extract subgraphs of relevant knowledge from input sentences. |
| Outcome: | The proposed algorithm achieves better quality-diversity tradeoff than a large language model with 100 times the number of parameters. |
Surface Form Competition: Why the Highest Probability Answer Isn’t Always Right (2021.emnlp-main)
Copied to clipboard
| Challenge: | Large language models have shown promising results in zero-shot settings due to surface form competition . since probability mass is finite, this lowers the probability of the correct answer . |
| Approach: | They propose a scoring function that compensates for surface form competition by reweighing each option according to its a priori likelihood. |
| Outcome: | The proposed scoring function achieves consistent gains in zero-shot over calibrated and uncalibrated scoring functions on all GPT-2 and GPT-3 models on a variety of multiple choice datasets. |
BottleHumor: Self-Informed Humor Explanation using the Information Bottleneck Principle (2025.findings-acl)
Copied to clipboard
| Challenge: | Humor is an effective communication tool that can manifest in various forms, including puns, exaggerated facial expressions, absurd behaviors, and incongruities. |
| Approach: | They propose a method that elicits relevant world knowledge from vision and language models and refines it to generate an explanation of the humor in an unsupervised manner. |
| Outcome: | The proposed method can be adapted for additional tasks that can benefit from eliciting and conditioning on relevant world knowledge. |
Stance Reasoner: Zero-Shot Stance Detection on Social Media with Explicit Reasoning (2024.lrec-main)
Copied to clipboard
| Challenge: | Stance Reasoner is a model for zero-shot stance detection on social media platforms that can be used to extract opinions from opinionated content. |
| Approach: | They propose a method that leverages explicit reasoning over background knowledge to guide the model’s inference about the document’s stance on a target. |
| Outcome: | The proposed model outperforms the current state-of-the-art models on 3 Twitter datasets, including fully supervised models. |
Paraphrasing vs Coreferring: Two Sides of the Same Coin (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Lexical resources such as WordNet capture synonyms and hypernyms, as well as antonyms, which can be used to refer to the same event when the arguments are reversed. |
| Approach: | They used annotations from an event coreference dataset as distant supervision to re-score heuristically-extracted predicate paraphrases. |
| Outcome: | The proposed model improved modestly but consistently in the two tasks. |