Papers by Anna Rogers

24 papers
Changing the World by Changing the Data (2021.acl-long)

Copied to clipboard

Challenge: a new paper argues that data curation is already happening, and it is changing the world . social biases and spurious patterns are attracting more attention in NLP models .
Approach: They argue that data curation is already happening and will be happening . they argue that social biases and spurious patterns are the main problems .
Outcome: a new paper argues that data curation is already and will be happening, and it is changing the world.
What Can We Do to Improve Peer Review in NLP? (2020.findings-emnlp)

Copied to clipboard

Challenge: Traditionally, peer review is expected to act as a filter for high-quality, impactful work, but this does not hold in practice.
Approach: They argue that peer review is becoming increasingly spurious and that it is a problem for NLP . they propose a reproducibility checklist at EMNLP 2020 that could be used to ensure that papers are reproducible.
Outcome: The reproducibility checklist at EMNLP 2020 is the first step in that direction.
RuSentiment: An Enriched Sentiment Analysis Dataset for Social Media in Russian (C18-1)

Copied to clipboard

Challenge: RuSentiment is currently the largest in its class for Russian, with 31,185 posts annotated with Fleiss’ kappa of 0.58 (3 annotations per post).
Approach: They propose to use RuSentiment to annotate social media posts in Russian with a kappa of 0.58 and a set of annotation guidelines that are extensible to other languages.
Outcome: The proposed dataset is the largest in its class for Russian, with 31,185 posts annotated with Fleiss’ kappa of 0.58 (3 annotations per post).
The ROOTS Search Tool: Data Transparency for LLMs (2023.acl-demo)

Copied to clipboard

Challenge: a 1.6TB multilingual text corpus is currently the largest language model . large language models are ubiquitous in modern NLP, used directly to generate text and as building blocks in downstream applications.
Approach: They propose a search engine for the 1.6TB multilingual ROOTS corpus offering both fuzzy and exact search capabilities.
Outcome: The ROOTS Search Tool is an open-source search engine for the 1.6TB multilingual ROOTs corpus.
Program Chairs’ Report on Peer Review at ACL 2023 (2023.acl-long)

Copied to clipboard

Challenge: ACL'23 makes its peer review report public and an official part of the conference proceedings.
Approach: They present an analysis of the factors affecting peer review and identify the most problematic issues that the authors complained about.
Outcome: The authors identified the most problematic issues and provided suggestions for the future chairs.
DECAF: A Dynamically Extensible Corpus Analysis Framework (2025.acl-demo)

Copied to clipboard

Challenge: DeCAF is an open-source Python library that enables the analysis and filtering of linguistically-annotated datasets down to the character level.
Approach: They propose a framework that enables the analysis and filtering of linguistically-annotated datasets down to the character level.
Outcome: The proposed framework analyzes a parsed version of the 115M-word BabyLM corpus and generates highly controlled and reproducible experimental settings targeting specific research questions.
Reviewing Natural Language Processing Research (2021.eacl-tutorials)

Copied to clipboard

Challenge: a tutorial on reviewing is a useful tool for researchers who are new to the field of NLP.
Approach: this tutorial provides an opportunity to learn the basics of reviewing . more experienced researchers might find this tutorial interesting to revise their reviewing procedure.
Outcome: This tutorial teaches researchers how to revise their reviewing procedure .
What’s in Your Embedding, And How It Predicts Task Performance (C18-1)

Copied to clipboard

Challenge: Attempts to find a single technique for general-purpose intrinsic evaluation of word embeddings have so far not been successful.
Approach: They propose a method that quantifies interpretable characteristics of word vector neighborhoods and shows how they correlate with performance on 14 extrinsic and intrinsic task datasets.
Outcome: The proposed approach enables multi-faceted evaluation, parameter search, and generally – a more principled, hypothesis-driven approach to development of distributional semantic representations.
What Factors Should Paper-Reviewer Assignments Rely On? Community Perspectives on Issues and Ideals in Conference Peer-Review (2022.naacl-main)

Copied to clipboard

Challenge: a survey of the NLP community shows that paper-reviewer matching is a problem . authors lose valuable time and opportunities by writing reviews that are arbitrarily low .
Approach: They propose to use paper-reviewer matching to improve peer review . they identify common issues and perspectives on what factors should be considered .
Outcome: The proposed method improves the quality of peer review and improves interpretable peer review assignments.
Revealing the Dark Secrets of BERT (D19-1)

Copied to clipboard

Challenge: Existing models of BERT-based learning systems are lacking specific mechanisms that contribute to its success.
Approach: They propose to use GLUE tasks to analyze the interpretation of self-attention, which is one of the underlying components of BERT.
Outcome: The proposed model outperforms the regular model on GLUE tasks by disabling attention in certain heads.
Adversarial Decomposition of Text Representation (N19-1)

Copied to clipboard

Challenge: a new method for adversarial decomposition of text representations is proposed . it is capable of fine-grained controlled change of different aspects of the input sentence .
Approach: They propose a method for adversarial decomposition of text representation . they use vectors responsible for a specific aspect of the input sentence .
Outcome: The proposed method outperforms the embeddings of a regular autoencoder on paraphrase detection tasks.
On the Interaction of Belief Bias and Explanations (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods to evaluate explainability fail to account for belief biases affecting human performance . previous studies have shown that neural models can make confident predictions relying on artifacts .
Approach: They propose to account for belief bias in explainability by using models of varying quality and adversarial examples.
Outcome: The proposed methods show that results change when using models of varying quality and adversarial examples.
Outlier Dimensions that Disrupt Transformers are Driven by Frequency (2022.findings-emnlp)

Copied to clipboard

Challenge: Disabling only 48 out of 110M parameters in BERT-base drops its performance by nearly 30% on MNLI.
Approach: They find that disabling only 48 out of 110M parameters in BERT-base drops its performance by nearly 30% on MNLI.
Outcome: The proposed model outlier phenomenon is associated with the frequency of encoded tokens in pre-training data.
A guide to the dataset explosion in QA, NLI, and commonsense reasoning (2020.coling-tutorials)

Copied to clipboard

Challenge: a tutorial aims to provide an up-to-date guide to the recent datasets . the target audience is the NLP practitioners who are lost in dozens of the recent data sets.
Approach: This tutorial provides an up-to-date guide to the recent datasets . it surveys old and new methodological issues with dataset construction .
Outcome: This tutorial aims to provide an up-to-date guide to the recent datasets . it surveys the old and new methodological issues with dataset construction .
Research Community Perspectives on “Intelligence” and Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Despite the widespread use of ‘artificial intelligence’ (AI) framing in NLP research, it is not clear what researchers mean by ”intelligence”.
Approach: They propose to use the term "AI" to describe the perception of a system as intelligent, but note that it is not accepted by the majority of respondents.
Outcome: The results suggest that the perception of the current NLP systems as 'intelligent' is a minority position (29%).
Calls to Action on Social Media: Detection, Social Impact, and Censorship Potential (D19-50)

Copied to clipboard

Challenge: Calls to action are effective means of mobilization in social networks, but their potential for censorship and predicting offline protest events has not yet been evaluated.
Approach: They examine the possibility of their automatic detection on historical data from the 2011-2013 protests in Bolotnaya, Russia.
Outcome: The political calls to action can be annotated and detected with relatively high accuracy and have a moderate positive correlation with actual rally attendance.
When BERT Plays the Lottery, All Tickets Are Winning (2020.emnlp-main)

Copied to clipboard

Challenge: Large Transformer-based models are reduced to a smaller number of self-attention heads and layers.
Approach: They propose to prune BERT self-attention heads and layers to find subnetworks with comparable performance . they also extend this technique to multi-layer perceptrons to find out if they are unstable .
Outcome: The proposed models are able to achieve 90% of full model performance with structured pruning and similar-sized subnetworks sampled from the rest of the model perform worse.
Code Like Humans: A Multi-Agent Solution for Medical Coding (2025.findings-emnlp)

Copied to clipboard

Challenge: In medical coding, experts map unstructured clinical notes to alphanumeric codes for diagnoses and procedures.
Approach: They introduce ‘Code Like Humans’: a new agentic framework for medical coding with large language models that implements official coding guidelines for human experts.
Outcome: The proposed framework implements official coding guidelines for human experts and can support the full ICD-10 coding system (+70K labels).
Machine Reading, Fast and Slow: When Do Models “Understand” Language? (2022.coling-1)

Copied to clipboard

Challenge: Existing models of reading comprehension score highly on NLU benchmarks, but they are often 'read fast', i.e. rely on shallow patterns.
Approach: They propose a definition for the reasoning steps expected from a system that would be 'reading slowly' they compare that behavior with five models of the BERT family of various sizes, observed through saliency scores and counterfactual explanations.
Outcome: The proposed model is compared with five models of the BERT family of various sizes, and compared using saliency scores and counterfactual explanations.
BERT Busters: Outlier Dimensions that Disrupt Transformers (2021.findings-acl)

Copied to clipboard

Challenge: Existing studies show that pre-trained Transformers are remarkably robust to pruning.
Approach: They show that pre-trained Transformer encoders are surprisingly fragile to pruning . they show that disabling them significantly degrades both the MLM loss and the downstream task performance.
Outcome: The results show that the removal of features in pre-trained transformers significantly degrades both the MLM loss and the downstream task performance.
NarrativeTime: Dense Temporal Annotation on a Timeline (2024.lrec-main)

Copied to clipboard

Challenge: e.g. TimeBank contains 1-5% of all possible tlinks, and this information is underspecified in the text.
Approach: They propose a timeline-based framework that achieves full coverage of all possible TLINKs.
Outcome: The proposed framework achieves full coverage of all possible TLINKs in a text.
A Primer in BERTology: What We Know About How BERT Works (2020.tacl-1)

Copied to clipboard

Challenge: a new study examines the current state of knowledge about the BERT model . the model is a stack of transformer encoder layers that are based on multiple self-attention ''heads''
Approach: They present a survey of over 150 studies of the popular Transformer-based model BERT . they discuss the current state of knowledge about how BERT works and how it is represented .
Outcome: The proposed model is based on the Transformer-based model with state-of-the-art results . the proposed model has little cognitive motivation and is too small to perform ablation studies .
‘Just What do You Think You’re Doing, Dave?’ A Checklist for Responsible Data Use in NLP (2021.findings-emnlp)

Copied to clipboard

Challenge: a key part of the NLP ethics movement is responsible use of data, but what that means is unclear . a proposed checklist for responsible data (re-)use could standardise peer review of submissions .
Approach: They propose a checklist for responsible data use that could standardise peer review . they propose implementing a standard for data (re-)use across NLP conferences .
Outcome: The proposed checklist would standardise peer review of submissions and enable more in-depth view of published research across the community.
AInterviewer: A Platform for Designing and Conducting AI-led Qualitative Interviews (2026.acl-demo)

Copied to clipboard

Challenge: Existing systems rely on proprietary LLMs, which compromise reproducibility and data security.
Approach: They propose a platform that combines controlled question administration of survey software with the flexibility of LLMs.
Outcome: AInterviewer combines controlled question administration of survey software with flexibility of LLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations