Papers by Maria Barrett

10 papers
Adversarial Removal of Demographic Attributes Revisited (D19-1)

Copied to clipboard

Challenge: Several approaches have been proposed to learn classifiers that are invariant (unbiased with respect) to protected attributes.
Approach: They propose to use a diagnostic classifier trained on a held-out subsample to find protected attributes for mention detection at above-chance levels.
Outcome: The proposed classifier generalizes poorly to new in-domain and new domains, suggesting it relies on correlations specific to their particular data sample.
The Copenhagen Corpus of Eye Tracking Recordings from Natural Reading of Danish Texts (2022.lrec-1)

Copied to clipboard

Challenge: Corpora of eye movements during reading of contextualized running text is a way of making such records available for natural language processing.
Approach: They present CopCo, the first eye tracking corpus of its kind for the Danish language.
Outcome: The Copenhagen corpus of eye tracking recordings from natural reading of Danish texts is the first of its kind for the Danish language.
Spurious Correlations in Cross-Topic Argument Mining (2021.starsem-1)

Copied to clipboard

Challenge: Recent work in cross-topic argument mining attempts to learn models that generalise across topics rather than relying on within-topic spurious correlations.
Approach: They propose to use linear approximations of decision boundaries and manual feature grouping to learn models that generalise across topics rather than relying on within-topic spurious correlations.
Outcome: The proposed model generalise across topics rather than relying on spurious correlations.
DaNE: A Named Entity Resource for Danish (2020.lrec-1)

Copied to clipboard

Challenge: a named entity annotation for the Danish Universal Dependencies treebank is the largest publicly available named entity gold annotation.
Approach: They propose a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme DaNE.
Outcome: The proposed annotations improve Danish named entity recognition over a recent cross-lingual approach and over norwegian training set.
Type B Reflexivization as an Unambiguous Testbed for Multilingual Multi-Task Gender Bias (2020.emnlp-main)

Copied to clipboard

Challenge: English challenge datasets highlight gender-ambiguous occurrences of ‘doctor’ as male doctors, but they are not useful for other languages.
Approach: They propose to build multi-task challenge datasets for detecting gender bias that lead to unambiguously wrong model predictions for languages with type B reflexivization.
Outcome: The proposed dataset can detect gender bias in languages with type B reflexivization and spans four languages and four NLP tasks.
Can Humans Identify Domains? (2024.lrec-main)

Copied to clipboard

Challenge: Textual domain is a crucial property within the Natural Language Processing community due to its effects on downstream model performance.
Approach: They examine the level of human disagreement and the relative difficulty of each annotation task by training classifiers to perform the same task.
Outcome: The authors show that human proficiency in identifying related intrinsic textual properties is low and that disagreements are high.
Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns (2026.findings-acl)

Copied to clipboard

Challenge: Prior work has shown that large language models can successfully persuade humans and amplify persuasive language.
Approach: They propose a framework for evaluating how persuasive language generation is affected by recipient gender, sender intent, or output language.
Outcome: The proposed framework varies persuasive language when the recipient gender is specified or when the sender intent is specified.
The Sensitivity of Language Models and Humans to Winograd Schema Perturbations (2020.acl-main)

Copied to clipboard

Challenge: Large-scale pre-trained language models are driving recent improvements in perfromance on the Winograd Schema Challenge . a diagnostic dataset shows that these models are sensitive to linguistic perturbations that minimally affect human understanding .
Approach: They propose to use a dataset to test pre-trained language models for the Winograd Schema Challenge . they show that these models are sensitive to linguistic perturbations that minimally affect human understanding .
Outcome: The proposed models are sensitive to linguistic perturbations that minimally affect human understanding.
Unsupervised Induction of Linguistic Categories with Records of Reading, Speaking, and Writing (N18-1)

Copied to clipboard

Challenge: a few researchers have shown that data traces from human processing can be used to improve NLP models.
Approach: They propose to use data readily available for most languages to improve unsupervised induction . they find that english unsupervised POS induction achieves an error reduction of 1.5% .
Outcome: The proposed model improves on Ontonotes domains with a word embeddings.
Native Language Prediction from Gaze: a Reproducibility Study (2023.acl-srw)

Copied to clipboard

Challenge: Existing studies have shown that the linguistic properties of a speaker’s native language affect the cognitive processing of other languages.
Approach: They found that the correlation between eye movements and native language similarity may be more complex than the original study found.
Outcome: The proposed model shows that the correlation between eye movements and native language similarity may be more complex than the original study.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations