Papers by Maria Barrett
Adversarial Removal of Demographic Attributes Revisited (D19-1)
Copied to clipboard
| Challenge: | Several approaches have been proposed to learn classifiers that are invariant (unbiased with respect) to protected attributes. |
| Approach: | They propose to use a diagnostic classifier trained on a held-out subsample to find protected attributes for mention detection at above-chance levels. |
| Outcome: | The proposed classifier generalizes poorly to new in-domain and new domains, suggesting it relies on correlations specific to their particular data sample. |
The Copenhagen Corpus of Eye Tracking Recordings from Natural Reading of Danish Texts (2022.lrec-1)
Copied to clipboard
| Challenge: | Corpora of eye movements during reading of contextualized running text is a way of making such records available for natural language processing. |
| Approach: | They present CopCo, the first eye tracking corpus of its kind for the Danish language. |
| Outcome: | The Copenhagen corpus of eye tracking recordings from natural reading of Danish texts is the first of its kind for the Danish language. |
Spurious Correlations in Cross-Topic Argument Mining (2021.starsem-1)
Copied to clipboard
| Challenge: | Recent work in cross-topic argument mining attempts to learn models that generalise across topics rather than relying on within-topic spurious correlations. |
| Approach: | They propose to use linear approximations of decision boundaries and manual feature grouping to learn models that generalise across topics rather than relying on within-topic spurious correlations. |
| Outcome: | The proposed model generalise across topics rather than relying on spurious correlations. |
DaNE: A Named Entity Resource for Danish (2020.lrec-1)
Copied to clipboard
Rasmus Hvingelby, Amalie Brogaard Pauli, Maria Barrett, Christina Rosted, Lasse Malm Lidegaard, Anders Søgaard
| Challenge: | a named entity annotation for the Danish Universal Dependencies treebank is the largest publicly available named entity gold annotation. |
| Approach: | They propose a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme DaNE. |
| Outcome: | The proposed annotations improve Danish named entity recognition over a recent cross-lingual approach and over norwegian training set. |
Type B Reflexivization as an Unambiguous Testbed for Multilingual Multi-Task Gender Bias (2020.emnlp-main)
Copied to clipboard
| Challenge: | English challenge datasets highlight gender-ambiguous occurrences of ‘doctor’ as male doctors, but they are not useful for other languages. |
| Approach: | They propose to build multi-task challenge datasets for detecting gender bias that lead to unambiguously wrong model predictions for languages with type B reflexivization. |
| Outcome: | The proposed dataset can detect gender bias in languages with type B reflexivization and spans four languages and four NLP tasks. |
Can Humans Identify Domains? (2024.lrec-main)
Copied to clipboard
Maria Barrett, Max Müller-Eberstein, Elisa Bassignana, Amalie Brogaard Pauli, Mike Zhang, Rob van der Goot
| Challenge: | Textual domain is a crucial property within the Natural Language Processing community due to its effects on downstream model performance. |
| Approach: | They examine the level of human disagreement and the relative difficulty of each annotation task by training classifiers to perform the same task. |
| Outcome: | The authors show that human proficiency in identifying related intrinsic textual properties is low and that disagreements are high. |
Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns (2026.findings-acl)
Copied to clipboard
| Challenge: | Prior work has shown that large language models can successfully persuade humans and amplify persuasive language. |
| Approach: | They propose a framework for evaluating how persuasive language generation is affected by recipient gender, sender intent, or output language. |
| Outcome: | The proposed framework varies persuasive language when the recipient gender is specified or when the sender intent is specified. |
The Sensitivity of Language Models and Humans to Winograd Schema Perturbations (2020.acl-main)
Copied to clipboard
| Challenge: | Large-scale pre-trained language models are driving recent improvements in perfromance on the Winograd Schema Challenge . a diagnostic dataset shows that these models are sensitive to linguistic perturbations that minimally affect human understanding . |
| Approach: | They propose to use a dataset to test pre-trained language models for the Winograd Schema Challenge . they show that these models are sensitive to linguistic perturbations that minimally affect human understanding . |
| Outcome: | The proposed models are sensitive to linguistic perturbations that minimally affect human understanding. |
Unsupervised Induction of Linguistic Categories with Records of Reading, Speaking, and Writing (N18-1)
Copied to clipboard
| Challenge: | a few researchers have shown that data traces from human processing can be used to improve NLP models. |
| Approach: | They propose to use data readily available for most languages to improve unsupervised induction . they find that english unsupervised POS induction achieves an error reduction of 1.5% . |
| Outcome: | The proposed model improves on Ontonotes domains with a word embeddings. |
Native Language Prediction from Gaze: a Reproducibility Study (2023.acl-srw)
Copied to clipboard
| Challenge: | Existing studies have shown that the linguistic properties of a speaker’s native language affect the cognitive processing of other languages. |
| Approach: | They found that the correlation between eye movements and native language similarity may be more complex than the original study found. |
| Outcome: | The proposed model shows that the correlation between eye movements and native language similarity may be more complex than the original study. |