Papers by Wesley Scivetti
Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not Meaning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent evidence suggests that language models with human-scale pretraining data may possess a similar generalization ability by generalizing from frequent to rare constructions. |
| Approach: | They construct a synthetic benchmark that targets syntactic and semantic properties of the English Let-Alone construction and compare it with a human-scale transformer language model. |
| Outcome: | The proposed model can generalize from frequent to rare constructions, but human-scale models do not make correct generalizations about Let-Alone’s meaning. |
GDTB: Genre Diverse Data for English Shallow Discourse Parsing across Modalities, Text Types, and Domains (2024.emnlp-main)
Copied to clipboard
Yang Janet Liu, Tatsuya Aoyama, Wesley Scivetti, Yilun Zhu, Shabnam Behzad, Lauren Levine, Jessica Lin, Devika Tiwari, Amir Zeldes
| Challenge: | Existing shallow discourse parsing systems focus on the Wall Street Journal corpus, but the data is limited to the news domain and is 35 years old. |
| Approach: | They propose to use the Wall Street Journal corpus as a benchmark for PDTB-style shallow discourse parsing. |
| Outcome: | The proposed dataset is compatible with PDTB, but suffers from degradation out-of-domain. |
Multilingual Supervision Improves Semantic Disambiguation of Adpositions (2025.coling-main)
Copied to clipboard
| Challenge: | a corpus-based cross-linguistic investigation into the lexical semantics of adpositions is conducted . a significant amount of ambiguity and flexibility in their meanings are present in a variety of languages . |
| Approach: | They conduct a corpus-based corpus analysis of adpositions using SNACS . they find distributional differences in a language's adequacy and disambiguation performance . |
| Outcome: | The proposed framework is suited for analyzing adpositions across languages . it provides a framework for a wide-coverage corpus annotation of high-level senses . |