Papers by Lea Frermann
Fairness-aware Class Imbalanced Learning (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on class imbalance and mitigating bias have focused on the latter . a skewed class distribution hurts the performance of deep learning models, and is often referred to as "stereotyping" |
| Approach: | They propose to extend a margin-loss based approach to enforce fairness by using tweet sentiment and occupation classification to mitigate class imbalance and demographic bias. |
| Outcome: | The proposed methods help mitigate class imbalance and demographic biases through controlled experiments. |
FairLib: A Unified Framework for Assessing and Improving Fairness (2022.emnlp-demos)
Copied to clipboard
| Challenge: | Existing approaches to assess and improve model fairness have been inconsistent and inconsistent. |
| Approach: | They propose an open-source python library for assessing and improving model fairness. |
| Outcome: | The proposed framework can be used for natural language, images, and audio. |
Book QA: Stories of Challenges and Opportunities (D19-58)
Copied to clipboard
| Challenge: | Existing approaches to answer questions based on the full text of books are limited by their unique characteristics. |
| Approach: | They propose a system for answering questions based on the full text of books . they use a memory network to reason and predict an answer, and a novel question generator to improve generalization. |
| Outcome: | The proposed system improves on the recently published NarrativeQA corpus on Who questions . it shows that the proposed system is highly challenging and needs more research . |
Evaluating Debiasing Techniques for Intersectional Biases (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for debiasing protected attributes have been limited to binary attributes in isolation, however many corpora involve multiple such attributes, possibly with higher cardinality. |
| Approach: | They propose to evaluate a bias-constrained model which is new to NLP and an extension of the iterative nullspace projection technique which can handle multiple identities. |
| Outcome: | The proposed model is based on a new iterative nullspace projection technique which can handle multiple identities. |
Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation methods focus on single-round inference, but this view is problematic in real-world applications. |
| Approach: | They propose a framework that couples Steering Token Calibration with Semantic Alignment to ensure that LLMs are correctly aligned across gender, race, and sentiment. |
| Outcome: | The proposed framework outperforms baseline methods in achieving precise distributional control in attribute generation tasks. |
Extractive NarrativeQA with Heuristic Pre-Training (D19-58)
Copied to clipboard
| Challenge: | Automated question answering (QA) from text remains a challenge for humans . a striking gap exists between machine and human performance on NLP tasks . |
| Approach: | They propose a heuristic extractive version of a data set to solve the problem of answer extraction rather than generation. |
| Outcome: | The proposed model outperforms previous models on summary-level QA from full narratives and on the METEOR metric. |
Optimising Equal Opportunity Fairness in Model Training (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing methods to reduce bias have been shown to be effective over real-world datasets. |
| Approach: | They propose two new training objectives which directly optimise for the widely-used criterion of equal opportunity. |
| Outcome: | The proposed training objectives directly optimise for the widely-used criterion of equal opportunity while maintaining high performance over two classification tasks. |
Probing Power by Prompting: Harnessing Pre-trained Language Models for Power Connotation Framing (2023.eacl-main)
Copied to clipboard
| Challenge: | Using pre-trained language models, we investigated whether word choices can encode subtle connotative information about power differentials between involved entities. |
| Approach: | They propose a framework to disentangle connotation frames implied by the predicate from its arguments and the sentence structure and to quantify predicates. |
| Outcome: | The proposed framework improves power connotation prediction accuracy by fine-tuning pre-trained language models. |
Unsupervised Cross-Lingual Transfer of Structured Predictors without Source Data (2022.naacl-main)
Copied to clipboard
| Challenge: | Recent successes of NLP systems require large amounts of labelled data for structured prediction tasks. |
| Approach: | They propose a method for unsupervised transfer from multiple input models for structured prediction using a cross-lingual setup. |
| Outcome: | The proposed method produces less noisy labels for the distant supervision. |
WAX: A New Dataset for Word Association eXplanations (2022.aacl-main)
Copied to clipboard
| Challenge: | Word associations are among the most common paradigms to study the human mental lexicon. |
| Approach: | They present a large dataset of word associations with explanations and relation labels . they show that current language models struggle to capture the diversity of human associations . |
| Outcome: | The proposed model fails to capture the diversity of human associations, the authors show . they show that the model is a rich benchmark for commonsense modeling and generation. |
Systematic Evaluation of Predictive Fairness (2022.aacl-main)
Copied to clipboard
| Challenge: | Several methods have been proposed to mitigate bias in training on biased datasets. |
| Approach: | They propose to examine the effect of target class imbalance and stereotyping on model performance by analyzing binary classification, profession prediction and regression tasks. |
| Outcome: | The proposed methods show that data conditions have a strong influence on relative model performance. |
CIG: Measuring Conversational Information Gain in Deliberative Dialogues with Semantic Memory Dynamics (2026.acl-long)
Copied to clipboard
| Challenge: | Using a semantic memory, we score each utterance along three interpretable dimensions: Novelty, Relevance, and Implication Scope. |
| Approach: | They propose a framework for Conversational Information Gain that evaluates each utterance in terms of how it advances collective understanding of the target topic. |
| Outcome: | The proposed framework evaluates each utterance in terms of how it advances collective understanding of the target topic. |
Inducing Document Structure for Aspect-based Summarization (P19-1)
Copied to clipboard
| Challenge: | Abstractive summarization systems treat documents as unstructured and generate a single generic summary per document. |
| Approach: | They propose to incorporate document structure into automatic summarization systems . they induce latent document structure and abstractive summarizing objective . |
| Outcome: | The proposed model improves on topic-agnostic baselines and can produce abstractive and extractive aspect-based summaries. |
Conflicts, Villains, Resolutions: Towards models of Narrative Media Framing (2023.acl-long)
Copied to clipboard
| Challenge: | a growing body of work attempts to automatically detect media frames in the news or social media, but most adopts a topic-like view on frames, evading modelling the broader document-level narrative. |
| Approach: | They propose an annotation paradigm that breaks a complex annotation task into a series of simple binary questions. |
| Outcome: | The proposed method is both effective and transparent in its predictions. |
A Computational Acquisition Model for Multimodal Word Categorization (2022.naacl-main)
Copied to clipboard
| Challenge: | Recent advances in self-supervised modeling of text and images open new opportunities for computational models of child language acquisition. |
| Approach: | They propose a multimodal language acquisition model trained from image-caption pairs on naturalistic data using cross-modal self-supervision. |
| Outcome: | The proposed model learns word categories and object recognition abilities, the authors show . their model is trained from image-caption pairs on naturalistic data using cross-modal self-supervision . |
Human Interest Framing across Cultures: A Case Study on Climate Change (2025.coling-main)
Copied to clipboard
| Challenge: | Human Interest (HI) framing is a narrative strategy that injects news stories with a relatable, emotional angle and a human face to engage the audience. |
| Approach: | They perform a systematic analysis of HI stories to understand its role in climate change reporting in English-speaking countries from four continents. |
| Outcome: | The proposed approach has shown to capture and retain readership and enhance political engagement of the population. |
Not all ANIMALs are equal: metaphorical framing through source domains and semantic frames (2026.findings-acl)
Copied to clipboard
| Challenge: | a computational framework allows to derive discourse metaphors through their source domains and semantic frames. |
| Approach: | They propose a computational framework that allows to derive salient discourse metaphors through their source domains and semantic frames. |
| Outcome: | The proposed framework uncovers well-known source domains and reveals nuanced frame-level associations that distinguish how the issue is portrayed. |
A Large-Scale Multilingual Study of Visual Constraints on Linguistic Selection of Descriptions (2023.findings-eacl)
Copied to clipboard
| Challenge: | a multilingual study examines how vision constrains linguistic choice . we use existing annotations to investigate the effect of different visual conditions on numeral expressions in captions . |
| Approach: | They propose a method that leverages existing corpora of images with captions written by native speakers to constrain linguistic choice. |
| Outcome: | The proposed method covers four languages and five linguistic properties, including verb transitivity and use of numerals. |
Unsupervised Induction of Linguistic Categories with Records of Reading, Speaking, and Writing (N18-1)
Copied to clipboard
| Challenge: | a few researchers have shown that data traces from human processing can be used to improve NLP models. |
| Approach: | They propose to use data readily available for most languages to improve unsupervised induction . they find that english unsupervised POS induction achieves an error reduction of 1.5% . |
| Outcome: | The proposed model improves on Ontonotes domains with a word embeddings. |
WHoW: A Cross-domain Approach for Analysing Conversation Moderation (2025.naacl-long)
Copied to clipboard
| Challenge: | Using this framework, we annotated 5,657 sentences with human judges and 15,494 sentences with GPT-4o from two domains: TV debates and radio panel discussions. |
| Approach: | They propose an evaluation framework for analyzing the facilitation strategies of moderators across different domains/scenarios by examining their motives (Why), dialogue acts (How) and target speaker (Who). |
| Outcome: | The framework is generalisable across domains and reveals distinct modes of moderation: debate moderators emphasise coordination and facilitate interaction through questions and instructions, panel discussion moderator prioritize information provision and actively participate in discussions. |
Narrative Media Framing in Political Discourse (2025.findings-acl)
Copied to clipboard
| Challenge: | Narrative frames are a powerful way of conceptualizing and communicating complex ideas. |
| Approach: | They propose a framework which formalizes and operationalizes elements of narrative framing . they annotate news articles in the climate change domain and test their framework . |
| Outcome: | The proposed framework formalizes and operationalizes elements of narrative framing . it is applied to climate change crisis data, showing generalizability of the framework . |
Partners in Crime: Multi-view Sequential Inference for Movie Understanding (D19-1)
Copied to clipboard
| Challenge: | Existing multi-view learning approaches are tested in unsupervised setups, allowing for learning of representation for monolithic data points, not sequences. |
| Approach: | They propose a neural architecture paired with a novel objective for incremental inference that integrates multi-view information for sequence prediction problems. |
| Outcome: | The proposed model outperforms previous work and strong baselines on two crime cases and speaker type tagging tasks that contribute to movie understanding. |
Article and Comment Frames Shape the Quality of Online Comments (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent work has focused on predicting comment toxicity or quality, but it ignores audience reactions. |
| Approach: | They propose a frame-aware system to mitigate unhealthy discourse . they analysed 1M comments across 2.7K news articles . |
| Outcome: | The proposed system can mitigate unhealthy discourses by analyzing 1M comments across 2.7K news articles. |
Moderation Matters: Measuring Conversational Moderation Impact in English as a Second Language Group Discussion (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing tools for ESL assessment focus on writing skills and lack in support for dynamic spoken interactions. |
| Approach: | They propose an approach that integrates automatic ESL dialogue assessment and a framework that categorizes moderation strategies to assess conversational engagement and moderation effectiveness. |
| Outcome: | The proposed approach integrates automatic ESL dialogue assessment and categorizes moderation strategies. |
More than Votes? Voting and Language based Partisanship in the US Supreme Court (2023.findings-emnlp)
Copied to clipboard
| Challenge: | partisanship and ideology have been a key topic in legal studies of the US Supreme Court . most research quantifies partisan behavior based on voting behavior, and oral arguments have not been well studied for this purpose. |
| Approach: | They propose a framework for analyzing justices' oral arguments for partisan signals and how they align with voting patterns. |
| Outcome: | The proposed framework shows that the affiliated party of justices can be predicted reliably from their oral contributions. |
Screenplay Summarization Using Latent Narrative Structure (2020.acl-main)
Copied to clipboard
| Challenge: | Experimental results show that latent turning points improve summarization performance over general extractive summarizing models. |
| Approach: | They propose to explicitly incorporate the underlying structure of narratives into extractive summarization models by treating it as latent. |
| Outcome: | The proposed model improves on the CSI corpus of screenplays on a CSI episode . it shows that latent turning points correlate with important aspects of the document . |
Comparing Moral Values in Western English-speaking societies and LLMs with Word Associations (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are trained on extensive corpora to learn linguistic patterns, contextual nuances, and implicit elements of human values. |
| Approach: | They propose to use word associations as low-level underlying representations to obtain a more robust picture of LLMs’ moral reasoning. |
| Outcome: | The proposed method reveals detailed but systematic differences between LLMs and human associations. |
Does Representational Fairness Imply Empirical Fairness? (2022.findings-aacl)
Copied to clipboard
| Challenge: | Neural methods have been trained on datasets which embody cultural and societal stereotypes, captured in spurious correlations between target labels and protected attributes. |
| Approach: | They propose a debiasing method that encourages a latent space that separates instances based on target label, while mixing instances that share protected attributes. |
| Outcome: | The proposed method shows that representational fairness does not imply empirical fairness across methods. |
PPT: Parsimonious Parser Transfer for Unsupervised Cross-Lingual Adaptation (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing methods for cross-lingual transfer use implicit supervision to parse low-resource languages without explicit supervision. |
| Approach: | They propose a method for unsupervised cross-lingual transfer that uses their output as implicit supervision as part of self-training on unlabelled text in the target language. |
| Outcome: | The proposed method improves over state-of-the-art models on both distant and nearby languages, despite being conceptually simpler. |
Seeking Clozure: Robust Hypernym extraction from BERT with Anchored Prompts (2023.starsem-1)
Copied to clipboard
| Challenge: | Existing methods for extracting hypernym knowledge from large language models are unclear whether they fail due to a lack of knowledge or shortcomings. |
| Approach: | They propose to use pattern-based hypernym extraction as a diagnostic tool to examine hypernomy knowledge encoded in BERT. |
| Outcome: | The proposed method compares the results of two different methods on six English data sets and on challenge sets of rare and abstract concepts. |
Framing Unpacked: A Semi-Supervised Interpretable Multi-View Model of Media Frames (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing models for news analysis lack transparency in their predictions. |
| Approach: | They propose a semi-supervised model that embeds local information into news articles . it can be used to improve automatic news analysis, authors argue . |
| Outcome: | The proposed model outperforms previous models and can be used with unlabeled training data. |