Papers by Shuly Wintner
Speaker Information Can Guide Models to Better Inductive Biases: A Case Study On Predicting Code-Switching (2022.acl-long)
Copied to clipboard
| Challenge: | Prior approaches for predicting code-switching only consider shallow linguistic context. |
| Approach: | They hypothesize that enriching models with speaker information can guide them to pick up on relevant inductive biases. |
| Outcome: | The proposed model improves on a speaker-driven task in English–Spanish bilingual dialogues by adding sociolinguistically-grounded speaker features as prepended prompts. |
Framing and Agenda-setting in Russian News: a Computational Analysis of Intricate Political Strategies (D18-1)
Copied to clipboard
| Challenge: | Amidst growing concern over media manipulation, NLP studies focus on overt strategies like censorship and “fake news”. |
| Approach: | They propose to use two concepts from political science literature to identify subtler media manipulation strategies . they propose to apply embedding-based methods to cross-lingually project English frames to Russian . |
| Outcome: | The proposed techniques can be applied to 13 years of the Russian newspaper Izvestia and show that they highlight U.S. moral failings and threats to the U.s. |
Machine Translation into Low-resource Language Varieties (2021.acl-short)
Copied to clipboard
| Challenge: | Current machine translation systems generate a "standard" target language, but many languages have multiple varieties that are different from the standard language. |
| Approach: | They propose a framework to rapidly adapt machine translation systems to generate different target varieties . they propose to use no parallel data to generate languages close to, but different from, the standard target language . |
| Outcome: | The proposed model improves on a system that generates Ukrainian and Belarusian in two languages with no parallel data. |
Topics to Avoid: Demoting Latent Confounds in Text Classification (D19-1)
Copied to clipboard
| Challenge: | Despite impressive performance on many text classification tasks, deep neural networks tend to learn frequent superficial patterns that are specific to the training data and do not always generalize well. |
| Approach: | They propose a method that represents latent topical confounds and a model which “unlearns” confounding features by predicting both the label of the input text and the confound. |
| Outcome: | The proposed model generalizes better and learns features indicative of the writing style rather than the content. |
Native Language Identification with User Generated Content (D18-1)
Copied to clipboard
| Challenge: | Using both linguistically-motivated features and the characteristics of the social media outlet, we obtain high accuracy on this challenging task. |
| Approach: | They propose to use linguistically-motivated features and social media characteristics to obtain high accuracy on this task. |
| Outcome: | The proposed method is highly accurate on a social media content where authors are highly-fluent nonnative speakers. |
Predicting the Proficiency Level of Nonnative Hebrew Authors (2022.lrec-1)
Copied to clipboard
| Challenge: | a recent study shows that nonnative Hebrew learners can be accurately predicted from their essays . the proficiency level of nonnativ speakers is important for educational purposes . |
| Approach: | They propose to use feature-based classifiers to accurately predict the proficiency level of nonnative Hebrew learners. |
| Outcome: | The proposed classifiers can predict the proficiency level of nonnative Hebrew learners . the results are compared with human graders on a corpus of Hebrew essays . |
The Hebrew Essay Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Annotated corpus of argumentative essays authored by prospective higher-education students . corpus includes essays by native speakers and essays by non-native speakers . |
| Approach: | They propose to use an annotated corpus of Hebrew argumentative essays to analyze non-native language use. |
| Outcome: | The proposed corpus includes essays by native speakers and essays authored by non-native speakers with three different native languages. |