Papers by Emily Allaway
Event-Guided Denoising for Multilingual Relation Learning (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods for general purpose relation extraction use a fixed set of predetermined relations, but research has shifted to the identification of unseen relations in any language. |
| Approach: | They propose a method for collecting high quality relation training data for relation extraction from unlabeled text that achieves a near-recreation of their zero-shot and few-shot results at a fraction of the training cost. |
| Outcome: | The proposed method achieves comparable results to the current state-of-the-art when trained on a smaller multilingual encoder . |
Sequential Cross-Document Coreference Resolution (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing models for cross-document coreference resolution have been used for within-document entity coreference but have been relatively limited. |
| Approach: | They propose a model that extends the efficient sequential prediction paradigm for coreference resolution to cross-document settings and achieves competitive results for both entity and event coreference. |
| Outcome: | The proposed model achieves competitive results for entity and event coreference while minimizing error propagation in complex reasoning tasks. |
Beyond Denouncing Hate: Strategies for Countering Implied Biases and Stereotypes in Language (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Counterspeech, i.e. responses to counteract potential harms of hateful speech, has become an increasingly popular solution to address online hate speech without censorship risks of deletion-based content moderation. |
| Approach: | They draw from psychology and philosophy literature to craft six psychologically inspired strategies to challenge the underlying stereotypical implications of hateful language. |
| Outcome: | The strategies used in human- and machine-generated counterspeech datasets are convincing, whereas human-written counterspech uses less specific strategies compared to machine-produced counters. |
Does Putting a Linguist in the Loop Improve NLU Data Collection? (2021.findings-emnlp)
Copied to clipboard
Alicia Parrish, William Huang, Omar Agha, Soo-Hwan Lee, Nikita Nangia, Alexia Warstadt, Karmanya Aggarwal, Emily Allaway, Tal Linzen, Samuel R. Bowman
| Challenge: | Many datasets for training and evaluating natural language understanding (NLU) models contain systematic artifacts that are identified only after data collection is complete. |
| Approach: | They propose to have linguists identify artifacts and gaps in the data and communicate with non-expert crowdworkers to adjust task instructions and incentives. |
| Outcome: | The proposed protocol does not increase accuracy on out-of-domain test sets, and adds a chatroom does not. |
Seeded Hierarchical Clustering for Expert-Crafted Taxonomies (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Practitioners from many disciplines use expert-crafted taxonomies to make sense of large, unlabeled corpora. |
| Approach: | They propose a weakly supervised algorithm for seeded hierarchical clustering that fits unlabeled data to taxonomies using a small set of labeled examples. |
| Outcome: | The proposed algorithm outperforms baselines on three real-world datasets. |
Evaluating Defeasible Reasoning in LLMs with DEFREASING (2025.naacl-long)
Copied to clipboard
| Challenge: | Defeasible inferences are highly plausible but can be impacted by new information. |
| Approach: | They construct a dataset to evaluate defeasible reasoning about property inheritance . they use generics to represent the inheritance rules because their semantics include exceptions . |
| Outcome: | The proposed model performs poorly across all pattern types and achieves 0.64 F 1 . the best performing model only achieves F 1 and the model is not well tuned . |
Zero-Shot Stance Detection: A Dataset and Model using Generalized Topic Representations (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for stance detection are topic-specific and cross-target stance. |
| Approach: | They propose a new dataset for zero-shot stance detection that captures a wider range of topics and lexical variation than in previous datasets. |
| Outcome: | The proposed model improves performance on a number of challenging linguistic phenomena. |
Human Rationales as Attribution Priors for Explainable Stance Detection (2021.emnlp-main)
Copied to clipboard
| Challenge: | In this work, we present a method for imparting human-like rationalization to a stance detection model using crowdsourced annotations on a small fraction of the training data. |
| Approach: | They propose a method for imparting human-like rationalization to a stance detection model using crowdsourced annotations on a small fraction of the training data. |
| Outcome: | The proposed method improves the reasoning of a state-of-the-art classifier in a data-scarce setting at no cost in predictive performance. |
Analyzing LLM Instruction Optimization for Tabular Fact Verification (2026.findings-eacl)
Copied to clipboard
Xiaotang Du, Giwon Hong, Wai-Chung Kwan, Rohit Saxena, Ivan Titov, Pasquale Minervini, Emily Allaway
| Challenge: | evaluating instruction optimization for tabular fact verification is a key challenge for reliable NLP systems. |
| Approach: | They compare instruction optimization for tabular fact verification with a framework based on DSPy . they find that instruction optimization consistently improves verification accuracy . |
| Outcome: | The proposed method improves verification accuracy across four benchmarks and three model families. |
Penguins Don’t Fly: Reasoning about Generics through Instantiations and Exceptions (2023.eacl-main)
Copied to clipboard
| Challenge: | Generics express generalizations about the world that are not universally true . commonsense knowledge bases encode some generic knowledge but rarely enumerate exceptions . |
| Approach: | They propose a framework informed by linguistic theory to generate exemplars for generics . they generate 19k exemplar cases for 650 generics and show they outperform a strong baseline . |
| Outcome: | The proposed framework outperforms a baseline framework by 12.8 precision points. |
A Unified Feature Representation for Lexical Connotations (2021.eacl-main)
Copied to clipboard
| Challenge: | ideological attitudes and stance are often expressed through subtle meanings of words and phrases. |
| Approach: | They propose a method for lexical representations that capture connotations within the embedding space . they define six new fine-grained connotation aspects for nouns and adjectives . |
| Outcome: | The proposed method improves stance detection when data is limited. |
Mitigating Covertly Unsafe Text within Natural Language Systems (2022.findings-emnlp)
Copied to clipboard
Alex Mei, Anisha Kabir, Sharon Levy, Melanie Subbiah, Emily Allaway, John Judge, Desmond Patton, Bruce Bimber, Kathleen McKeown, William Yang Wang
| Challenge: | Existing studies on text safety have focused on overtly unsafe, covertly, or indirectly unsafe statements. |
| Approach: | They propose a method to identify physical harm-causing statements as overtly, covertly or indirectly unsafe and a solution to mitigate the generation of such statements. |
| Outcome: | The proposed methods identify the type of unsafe language that can cause physical harm and identify mitigation strategies to inspire future researchers to tackle this challenging problem. |
VISaGE: Understanding Visual Generics and Exceptions (2025.emnlp-main)
Copied to clipboard
| Challenge: | atypical evaluation instances disrupt incontext instance understanding and in-weight conceptual knowledge. |
| Approach: | They propose to use a dataset to analyze atypical visual and textual images to test their models. |
| Outcome: | The proposed model is based on a dataset consisting of typical and exceptional images. |
Event2Mind: Commonsense Inference on Events, Intents, and Reactions (P18-1)
Copied to clipboard
| Challenge: | Using a crowdsourced corpus of 25,000 event phrases, we construct a new task that uses commonsense reasoning to reason about the likely intents and reactions of the event participants. |
| Approach: | They construct a crowdsourced corpus of 25,000 event phrases and use them to construct 'commonsense inference' they demonstrate that neural encoder-decoder models can compose embedding representations of previously unseen events and reason about the likely intents and reactions of the event participants. |
| Outcome: | The proposed task can be used to uncover implicit gender inequality in movie scripts. |
Adversarial Learning for Zero-Shot Stance Detection on Social Media (2021.naacl-main)
Copied to clipboard
| Challenge: | a new model for zero-shot stance detection on Twitter uses adversarial learning to generalize across topics . previous work on zero- shot stance detector on English social media focuses on cross-target stances . |
| Approach: | They propose a model that uses adversarial learning to generalize across topics on Twitter . their model achieves state-of-the-art performance on unseen test topics . |
| Outcome: | The proposed model achieves state-of-the-art performance on unseen topics with minimal computational costs. |
Generics are puzzling. Can language models find the missing piece? (2025.coling-main)
Copied to clipboard
| Challenge: | Generic sentences express generalisations about the world without explicit quantification . human biases in stereotypes can be observed in language models, authors say . |
| Approach: | They analyze generic sentences to determine their quantification and quantify their implicit quantifications using language models. |
| Outcome: | The proposed model shows that generics are more context-sensitive than determiner quantifiers and express weak generalisations. |
Generics are not quantificational: A new path from language models to semantic theory (2026.findings-acl)
Copied to clipboard
| Challenge: | Generic sentences express generalizations that tolerate exceptions without explicitly communicating information about quantities. |
| Approach: | They compare generics and quantificational sentences to find out what quantifiers are . they argue that generics are not quantificationals, contrary to dominant views . |
| Outcome: | The proposed model recovers many semantic facts about quantifiers and their "quantificational counterparts". |
SafeText: A Benchmark for Exploring Physical Safety in Language Models (2022.emnlp-main)
Copied to clipboard
Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia Chilton, Desmond Patton, Kathleen McKeown, William Yang Wang
| Challenge: | Existing models that generate unsafe text are susceptible to the dangers of unsafe text generation and are deemed unsafe. |
| Approach: | They use a dataset to empirically study commonsense physical safety across various models for text generation and reasoning tasks. |
| Outcome: | The proposed model can generate unsafe text and reject it, but the different harms that can occur do not receive equal attention, which may consequently downplay certain harms. |