Papers by Anna Rogers
Changing the World by Changing the Data (2021.acl-long)
Copied to clipboard
| Challenge: | a new paper argues that data curation is already happening, and it is changing the world . social biases and spurious patterns are attracting more attention in NLP models . |
| Approach: | They argue that data curation is already happening and will be happening . they argue that social biases and spurious patterns are the main problems . |
| Outcome: | a new paper argues that data curation is already and will be happening, and it is changing the world. |
What Can We Do to Improve Peer Review in NLP? (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Traditionally, peer review is expected to act as a filter for high-quality, impactful work, but this does not hold in practice. |
| Approach: | They argue that peer review is becoming increasingly spurious and that it is a problem for NLP . they propose a reproducibility checklist at EMNLP 2020 that could be used to ensure that papers are reproducible. |
| Outcome: | The reproducibility checklist at EMNLP 2020 is the first step in that direction. |
RuSentiment: An Enriched Sentiment Analysis Dataset for Social Media in Russian (C18-1)
Copied to clipboard
| Challenge: | RuSentiment is currently the largest in its class for Russian, with 31,185 posts annotated with Fleiss’ kappa of 0.58 (3 annotations per post). |
| Approach: | They propose to use RuSentiment to annotate social media posts in Russian with a kappa of 0.58 and a set of annotation guidelines that are extensible to other languages. |
| Outcome: | The proposed dataset is the largest in its class for Russian, with 31,185 posts annotated with Fleiss’ kappa of 0.58 (3 annotations per post). |
The ROOTS Search Tool: Data Transparency for LLMs (2023.acl-demo)
Copied to clipboard
Aleksandra Piktus, Christopher Akiki, Paulo Villegas, Hugo Laurençon, Gérard Dupont, Sasha Luccioni, Yacine Jernite, Anna Rogers
| Challenge: | a 1.6TB multilingual text corpus is currently the largest language model . large language models are ubiquitous in modern NLP, used directly to generate text and as building blocks in downstream applications. |
| Approach: | They propose a search engine for the 1.6TB multilingual ROOTS corpus offering both fuzzy and exact search capabilities. |
| Outcome: | The ROOTS Search Tool is an open-source search engine for the 1.6TB multilingual ROOTs corpus. |
Program Chairs’ Report on Peer Review at ACL 2023 (2023.acl-long)
Copied to clipboard
| Challenge: | ACL'23 makes its peer review report public and an official part of the conference proceedings. |
| Approach: | They present an analysis of the factors affecting peer review and identify the most problematic issues that the authors complained about. |
| Outcome: | The authors identified the most problematic issues and provided suggestions for the future chairs. |
DECAF: A Dynamically Extensible Corpus Analysis Framework (2025.acl-demo)
Copied to clipboard
| Challenge: | DeCAF is an open-source Python library that enables the analysis and filtering of linguistically-annotated datasets down to the character level. |
| Approach: | They propose a framework that enables the analysis and filtering of linguistically-annotated datasets down to the character level. |
| Outcome: | The proposed framework analyzes a parsed version of the 115M-word BabyLM corpus and generates highly controlled and reproducible experimental settings targeting specific research questions. |
Reviewing Natural Language Processing Research (2021.eacl-tutorials)
Copied to clipboard
| Challenge: | a tutorial on reviewing is a useful tool for researchers who are new to the field of NLP. |
| Approach: | this tutorial provides an opportunity to learn the basics of reviewing . more experienced researchers might find this tutorial interesting to revise their reviewing procedure. |
| Outcome: | This tutorial teaches researchers how to revise their reviewing procedure . |
What’s in Your Embedding, And How It Predicts Task Performance (C18-1)
Copied to clipboard
| Challenge: | Attempts to find a single technique for general-purpose intrinsic evaluation of word embeddings have so far not been successful. |
| Approach: | They propose a method that quantifies interpretable characteristics of word vector neighborhoods and shows how they correlate with performance on 14 extrinsic and intrinsic task datasets. |
| Outcome: | The proposed approach enables multi-faceted evaluation, parameter search, and generally – a more principled, hypothesis-driven approach to development of distributional semantic representations. |
What Factors Should Paper-Reviewer Assignments Rely On? Community Perspectives on Issues and Ideals in Conference Peer-Review (2022.naacl-main)
Copied to clipboard
| Challenge: | a survey of the NLP community shows that paper-reviewer matching is a problem . authors lose valuable time and opportunities by writing reviews that are arbitrarily low . |
| Approach: | They propose to use paper-reviewer matching to improve peer review . they identify common issues and perspectives on what factors should be considered . |
| Outcome: | The proposed method improves the quality of peer review and improves interpretable peer review assignments. |
Revealing the Dark Secrets of BERT (D19-1)
Copied to clipboard
| Challenge: | Existing models of BERT-based learning systems are lacking specific mechanisms that contribute to its success. |
| Approach: | They propose to use GLUE tasks to analyze the interpretation of self-attention, which is one of the underlying components of BERT. |
| Outcome: | The proposed model outperforms the regular model on GLUE tasks by disabling attention in certain heads. |
Adversarial Decomposition of Text Representation (N19-1)
Copied to clipboard
| Challenge: | a new method for adversarial decomposition of text representations is proposed . it is capable of fine-grained controlled change of different aspects of the input sentence . |
| Approach: | They propose a method for adversarial decomposition of text representation . they use vectors responsible for a specific aspect of the input sentence . |
| Outcome: | The proposed method outperforms the embeddings of a regular autoencoder on paraphrase detection tasks. |
On the Interaction of Belief Bias and Explanations (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to evaluate explainability fail to account for belief biases affecting human performance . previous studies have shown that neural models can make confident predictions relying on artifacts . |
| Approach: | They propose to account for belief bias in explainability by using models of varying quality and adversarial examples. |
| Outcome: | The proposed methods show that results change when using models of varying quality and adversarial examples. |
Outlier Dimensions that Disrupt Transformers are Driven by Frequency (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Disabling only 48 out of 110M parameters in BERT-base drops its performance by nearly 30% on MNLI. |
| Approach: | They find that disabling only 48 out of 110M parameters in BERT-base drops its performance by nearly 30% on MNLI. |
| Outcome: | The proposed model outlier phenomenon is associated with the frequency of encoded tokens in pre-training data. |
A guide to the dataset explosion in QA, NLI, and commonsense reasoning (2020.coling-tutorials)
Copied to clipboard
| Challenge: | a tutorial aims to provide an up-to-date guide to the recent datasets . the target audience is the NLP practitioners who are lost in dozens of the recent data sets. |
| Approach: | This tutorial provides an up-to-date guide to the recent datasets . it surveys old and new methodological issues with dataset construction . |
| Outcome: | This tutorial aims to provide an up-to-date guide to the recent datasets . it surveys the old and new methodological issues with dataset construction . |
Research Community Perspectives on “Intelligence” and Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Despite the widespread use of ‘artificial intelligence’ (AI) framing in NLP research, it is not clear what researchers mean by ”intelligence”. |
| Approach: | They propose to use the term "AI" to describe the perception of a system as intelligent, but note that it is not accepted by the majority of respondents. |
| Outcome: | The results suggest that the perception of the current NLP systems as 'intelligent' is a minority position (29%). |
Calls to Action on Social Media: Detection, Social Impact, and Censorship Potential (D19-50)
Copied to clipboard
| Challenge: | Calls to action are effective means of mobilization in social networks, but their potential for censorship and predicting offline protest events has not yet been evaluated. |
| Approach: | They examine the possibility of their automatic detection on historical data from the 2011-2013 protests in Bolotnaya, Russia. |
| Outcome: | The political calls to action can be annotated and detected with relatively high accuracy and have a moderate positive correlation with actual rally attendance. |
When BERT Plays the Lottery, All Tickets Are Winning (2020.emnlp-main)
Copied to clipboard
| Challenge: | Large Transformer-based models are reduced to a smaller number of self-attention heads and layers. |
| Approach: | They propose to prune BERT self-attention heads and layers to find subnetworks with comparable performance . they also extend this technique to multi-layer perceptrons to find out if they are unstable . |
| Outcome: | The proposed models are able to achieve 90% of full model performance with structured pruning and similar-sized subnetworks sampled from the rest of the model perform worse. |
Code Like Humans: A Multi-Agent Solution for Medical Coding (2025.findings-emnlp)
Copied to clipboard
Andreas Geert Motzfeldt, Joakim Edin, Casper L. Christensen, Christian Hardmeier, Lars Maaløe, Anna Rogers
| Challenge: | In medical coding, experts map unstructured clinical notes to alphanumeric codes for diagnoses and procedures. |
| Approach: | They introduce ‘Code Like Humans’: a new agentic framework for medical coding with large language models that implements official coding guidelines for human experts. |
| Outcome: | The proposed framework implements official coding guidelines for human experts and can support the full ICD-10 coding system (+70K labels). |
Machine Reading, Fast and Slow: When Do Models “Understand” Language? (2022.coling-1)
Copied to clipboard
| Challenge: | Existing models of reading comprehension score highly on NLU benchmarks, but they are often 'read fast', i.e. rely on shallow patterns. |
| Approach: | They propose a definition for the reasoning steps expected from a system that would be 'reading slowly' they compare that behavior with five models of the BERT family of various sizes, observed through saliency scores and counterfactual explanations. |
| Outcome: | The proposed model is compared with five models of the BERT family of various sizes, and compared using saliency scores and counterfactual explanations. |
BERT Busters: Outlier Dimensions that Disrupt Transformers (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing studies show that pre-trained Transformers are remarkably robust to pruning. |
| Approach: | They show that pre-trained Transformer encoders are surprisingly fragile to pruning . they show that disabling them significantly degrades both the MLM loss and the downstream task performance. |
| Outcome: | The results show that the removal of features in pre-trained transformers significantly degrades both the MLM loss and the downstream task performance. |
NarrativeTime: Dense Temporal Annotation on a Timeline (2024.lrec-main)
Copied to clipboard
| Challenge: | e.g. TimeBank contains 1-5% of all possible tlinks, and this information is underspecified in the text. |
| Approach: | They propose a timeline-based framework that achieves full coverage of all possible TLINKs. |
| Outcome: | The proposed framework achieves full coverage of all possible TLINKs in a text. |
A Primer in BERTology: What We Know About How BERT Works (2020.tacl-1)
Copied to clipboard
| Challenge: | a new study examines the current state of knowledge about the BERT model . the model is a stack of transformer encoder layers that are based on multiple self-attention ''heads'' |
| Approach: | They present a survey of over 150 studies of the popular Transformer-based model BERT . they discuss the current state of knowledge about how BERT works and how it is represented . |
| Outcome: | The proposed model is based on the Transformer-based model with state-of-the-art results . the proposed model has little cognitive motivation and is too small to perform ablation studies . |
‘Just What do You Think You’re Doing, Dave?’ A Checklist for Responsible Data Use in NLP (2021.findings-emnlp)
Copied to clipboard
| Challenge: | a key part of the NLP ethics movement is responsible use of data, but what that means is unclear . a proposed checklist for responsible data (re-)use could standardise peer review of submissions . |
| Approach: | They propose a checklist for responsible data use that could standardise peer review . they propose implementing a standard for data (re-)use across NLP conferences . |
| Outcome: | The proposed checklist would standardise peer review of submissions and enable more in-depth view of published research across the community. |
AInterviewer: A Platform for Designing and Conducting AI-led Qualitative Interviews (2026.acl-demo)
Copied to clipboard
| Challenge: | Existing systems rely on proprietary LLMs, which compromise reproducibility and data security. |
| Approach: | They propose a platform that combines controlled question administration of survey software with the flexibility of LLMs. |
| Outcome: | AInterviewer combines controlled question administration of survey software with flexibility of LLMs. |