Papers by Valentina Pyatkin
Asking It All: Generating Contextualized Questions for any Semantic Role (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to question generation require conditioning on existing answers in text . previous work required human-curated templates, limiting coverage and question fluency . |
| Approach: | They propose a task of role question generation that produces a prototype and revises it to be contextually appropriate for the passage. |
| Outcome: | The proposed model generates diverse and well-formed questions for a large, broad-coverage ontology of predicates and roles. |
QANom: Question-Answer driven SRL for Nominalizations (2020.coling-main)
Copied to clipboard
Ayal Klein, Jonathan Mamou, Valentina Pyatkin, Daniela Stepanov, Hangfeng He, Dan Roth, Luke Zettlemoyer, Ido Dagan
| Challenge: | Traditionally, SRL annotations focus on verbal predicates, but other types of predicate are frequent in natural language. |
| Approach: | They propose a semantic scheme for capturing predicate-argument relations for nominalizations, termed QANom, using crowdsourcing and QA-driven annotations. |
| Outcome: | The proposed scheme outperforms existing annotations and is useful for downstream tasks. |
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance (2026.tacl-1)
Copied to clipboard
Paul Röttger, Musashi Hinck, Valentin Hofmann, Kobi Hackenburg, Valentina Pyatkin, Faeze Brahman, Dirk Hovy
| Challenge: | Large language models are helping millions of users write texts about diverse issues . issue bias is where an LLM tends to present just one perspective on a given issue . |
| Approach: | They construct a set of 2.49m realistic English-language prompts to measure issue bias in LLM writing assistance using 3.9k templates and 212 political issues from real user interactions. |
| Outcome: | The proposed model aligns more with US Democrat than Republican voter opinion on a subset of issues. |
Promptly Predicting Structures: The Return of Inference (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing methods for structured prediction rely on large labeled datasets. Existing approaches for structured predictions require detailed annotation guidelines about the task, the label set, and the interactions between labels. |
| Approach: | They propose a framework for constructing zero- and few-shot linguistic structure predictors using structural constraints and combinatorial inferences. |
| Outcome: | The proposed framework can be extended to build zero- and few-shot label predictors on two structured prediction tasks and five datasets. |
The Possible, the Plausible, and the Desirable: Event-Based Modality Detection for Language Processing (2021.acl-long)
Copied to clipboard
| Challenge: | Existing studies restrict modal expressions to a closed syntactic class . modal sense labels are vastly different across different studies, lacking an accepted standard . |
| Approach: | They propose a task where modal expressions can be words of any syntactic class and sense labels are drawn from a comprehensive taxonomy which harmonizes the modal concepts contributed by the different studies. |
| Outcome: | The proposed task is based on the Georgetown Gradable Modal Expressions corpus . it detects and classifies fine-grained modal concepts and associates them with modified events . |
QASem Parsing: Text-to-text Modeling of QA-based Semantics (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing work suggests the appeals of incorporating explicit semantic representations into NLP . semi-structured natural language structures provide an intermediate meaning-capturing representation . |
| Approach: | They propose a semi-structured natural-language representation of textual information . they examine input and output linearization strategies and multitask learning . |
| Outcome: | The proposed model is based on pre-trained sequence-to-sequence language models . it is easy to use and can be used for downstream tasks that benefit from it . |
Draw Me a Flower: Processing and Grounding Abstraction in Natural Language (2022.tacl-1)
Copied to clipboard
| Challenge: | Abstraction is a core tenet of human cognition and communication. yet, interpreting and grounding abstraction expressed in natural language (NL) has not been systematically studied in NLP. |
| Approach: | They propose a 2D instruction-following game that elicits abstract instructions from 4k natural language instructions. |
| Outcome: | The proposed method elicits 4k natural language instructions rich with diverse types of abstractions and assesses neural models. |
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations (2023.findings-emnlp)
Copied to clipboard
Kavel Rao, Liwei Jiang, Valentina Pyatkin, Yuling Gu, Niket Tandon, Nouha Dziri, Faeze Brahman, Yejin Choi
| Challenge: | Moral or ethical judgments rely heavily on the contexts in which they occur . a student model that produces defeasible contexts with improved validity, diversity, and defasibility is superior to intermediate student models . |
| Approach: | a new study uses a student model to provide contextualizations that make an action morally acceptable . the model is based on a dataset of 115K defeasible moral actions rated highly by human annotators . |
| Outcome: | The proposed model outperforms all intermediate models in a high-quality dataset . the model is based on 1.2M entries of contextualizations and rationales for 115K moral actions . |
“You Are An Expert Linguistic Annotator”: Limits of LLMs as Analyzers of Abstract Meaning Representation (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) demonstrate proficiency and fluency in the use of language, but do they have the linguistic knowledge to serve as an expert linguistic annotator? |
| Approach: | They examine the successes and limitations of large language models using the Abstract Meaning Representation (AMR) parsing formalism. |
| Outcome: | The proposed models can reproduce the basic format of AMR, as well as some core event, argument, and modifier structure, but they have virtually no fully accurate parses. |
Design Choices for Crowdsourcing Implicit Discourse Relations: Revealing the Biases Introduced by Task Design (2023.tacl-1)
Copied to clipboard
| Challenge: | Disagreement in natural language annotation has been studied from a perspective of biases introduced by the annotators and the annotation frameworks. |
| Approach: | They propose to analyze task design bias in crowdsourced annotations where lay annotators are used to elicit interpretations. |
| Outcome: | The proposed methods can push annotators towards certain relations and some discourse relation senses can be better elicited with one or the other approach. |
Revisiting Sentence Union Generation as a Testbed for Text Consolidation (2023.findings-acl)
Copied to clipboard
| Challenge: | In order to acquire knowledge on a new subject, it is often necessary to consult multiple sources of written information. |
| Approach: | They propose to revisit the sentence union generation task as an effective well-defined testbed for assessing text consolidation capabilities. |
| Outcome: | The proposed evaluation protocol includes human and automatic evaluations. |
Superlatives in Context: Modeling the Implicit Semantics of Superlatives (2025.naacl-long)
Copied to clipboard
| Challenge: | a study of superlatives shows that the semantics of superlations in context can be challenging for contemporary models. |
| Approach: | They propose a unified account of superlative semantics which allows for a broad-coverage annotation schema. |
| Outcome: | The proposed schema allows for interpreting superlative expressions and their semantic interpretations. |
Is It JUST Semantics? A Case Study of Discourse Particle Understanding in LLMs (2025.findings-acl)
Copied to clipboard
William Berkeley Sheffield, Kanishka Misra, Valentina Pyatkin, Ashwini Deo, Kyle Mahowald, Junyi Jessy Li
| Challenge: | Discourse particles are crucial elements that subtly shape the meaning of text. |
| Approach: | They examine the capacity of linguists to distinguish fine-grained senses of English *just* . they find that they struggle to fully capture more subtle nuances of discourse particles . |
| Outcome: | The study shows that linguists struggle to capture subtle nuances of discourse particles. |
Design Choices in Crowdsourcing Discourse Relation Annotations: The Effect of Worker Selection and Training (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent methods have obtained promising results by extracting relation labels from participants . obtaining linguistic annotations from novice crowdworkers is difficult . crowdsourcing allows for fast and cost-effective collection of labelled data, but because tasks need to be intuitive, crowdworker cannot be asked to perform them. |
| Approach: | They propose to use a selection-only approach to obtain linguistic annotations from novices . current study shows that the method is cost- and time-intensive . |
| Outcome: | The current study shows that selection and training improves the agreement between workers and gold labels, but the method is cost- and time-intensive. |
Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback (2025.acl-long)
Copied to clipboard
Lester James Validad Miranda, Yizhong Wang, Yanai Elazar, Sachin Kumar, Valentina Pyatkin, Faeze Brahman, Noah A. Smith, Hannaneh Hajishirzi, Pradeep Dasigi
| Challenge: | Learning from human feedback has enabled the alignment of language models (LMs) with human preferences. |
| Approach: | They propose a Hybrid Preference routER that defers an annotation to either humans or LMs, achieving better annotation quality while reducing the cost of human-only annotation. |
| Outcome: | The proposed model achieves better annotation quality while reducing the cost of human-only annotation. |
RewardBench: Evaluating Reward Models for Language Modeling (2025.findings-naacl)
Copied to clipboard
Nathan Lambert, Valentina Pyatkin, Jacob Morrison, Lester James Validad Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A. Smith, Hannaneh Hajishirzi
| Challenge: | Evaluating reward models presents an opportunity to understand the opaque technologies used for alignment of language models. |
| Approach: | They present a benchmark dataset and code-base for evaluation of reward models . they use prompt-chosen-rejected trios to benchmark how they perform on queries . |
| Outcome: | The proposed dataset compares RMs with other models on a set of questions. |
QADiscourse - Discourse Relations as QA Pairs: Representation, Crowdsourcing and Baselines (2020.emnlp-main)
Copied to clipboard
| Challenge: | Discourse relations describe how two propositions relate to one another . annotating discourse relations requires expert annotators . |
| Approach: | They propose a new representation of discourse relations as question-and-answer pairs that crowd-sources wide-coverage data annotated with discourse relations. |
| Outcome: | The proposed representation of discourse relations as QA pairs allows crowd-sourcing wide-coverage datasets annotated with discourse relations. |
ClarifyDelphi: Reinforced Clarification Questions with Defeasibility Rewards for Social and Moral Situations (2023.acl-long)
Copied to clipboard
Valentina Pyatkin, Jena D. Hwang, Vivek Srikumar, Ximing Lu, Liwei Jiang, Yejin Choi, Chandra Bhagavatula
| Challenge: | Changing contexts can flip the moral judgment of an action. |
| Approach: | They propose an interactive system that learns to ask clarification questions to elicit salient contexts of a social or moral situation. |
| Outcome: | The proposed system generates more relevant, informative and defeasible questions compared to baselines. |