Papers by Lucie Flek
Mitigating Toxic Degeneration with Empathetic Data: Exploring the Relationship Between Toxicity and Empathy (2022.naacl-main)
Copied to clipboard
| Challenge: | Recent work on controllable text generation has shown promise in successfully altering such text attributes. |
| Approach: | They propose to use empathetic data to reduce the toxicity of generated text by strategically sampling data based on empathy scores. |
| Outcome: | The proposed model significantly reduces the size of fine-tuning data to 7.5-30k samples while making significant improvements over state-of-the-art toxicity mitigation. |
FACTOID: A New Dataset for Identifying Misinformation Spreaders and Political Bias (2022.lrec-1)
Copied to clipboard
| Challenge: | Proactively identifying misinformation spreaders is an important step towards mitigating the impact of fake news on our society. |
| Approach: | They propose a new reddit dataset for fake news spreader analysis, called FACTOID, which tracks political discussions on Reddit since the beginning of 2020. |
| Outcome: | The proposed dataset contains over 4K users with 3.4M posts and includes their credibility level (very low to very high) and political bias strength (extreme right to extreme left). |
Explainable Hallucination through Natural Language Inference Mapping (2025.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) often generate hallucinated content, making it crucial to identify and quantify inconsistencies in their outputs. |
| Approach: | They propose a framework that maps entailment and contradiction relations between inputs and outputs using a natural language inference model. |
| Outcome: | The proposed framework outperforms state-of-the-art methods by five percentage points while providing clear, interpretable explanations. |
Appraisal Framework for Clinical Empathy: A Novel Application to Breaking Bad News Conversations (2024.lrec-main)
Copied to clipboard
| Challenge: | Empathy is essential in healthcare communication. |
| Approach: | They propose an annotation approach that draws on well-established frameworks for clinical empathy and breaking bad news conversations for considering the dynamic dynamics of discourse relations. |
| Outcome: | The proposed model can be used to train models to detect causal relations involving empathy, a feature of systems that can provide feedback to medical professionals in training. |
Corpus Considerations for Annotator Modeling and Scaling (2024.naacl-long)
Copied to clipboard
| Challenge: | Recent trends in natural language processing and annotation tasks emphasize individual perspectives . annotator models that rely on a single ground truth may disregard valuable minority perspectives omissions . |
| Approach: | They propose a composite embedding approach to investigate annotator modeling techniques . they show that the commonly used user token model consistently outperforms more complex models . |
| Outcome: | The proposed model outperforms more complex models on a given dataset. |
Unifying Data Perspectivism and Personalization: An Application to Social Norms (2022.emnlp-main)
Copied to clipboard
| Challenge: | Obtaining a single ground truth is not possible or necessary for subjective tasks. |
| Approach: | They propose a set of personalization methods to model annotators and compare their effectiveness for predicting social norms. |
| Outcome: | The proposed model outperforms existing models and compares performance across subsets of social situations that vary by the closeness of the relationship between parties in conflict. |
Returning the N to NLP: Towards Contextually Personalized Classification Models (2020.acl-main)
Copied to clipboard
| Challenge: | a recent study shows that NLP models treat language as universal, but that it is based on sociolinguistic research. |
| Approach: | They propose to incorporate user-dependent, contextual personal and social aspects into neural NLP models by means of socially contextual personalization. |
| Outcome: | The proposed approach could be adapted to better personalize the language of users . it outlines a possible direction to incorporate these aspects into neural NLP models . |
Perspective Taking through Generating Responses to Conflict Situations (2024.findings-acl)
Copied to clipboard
| Challenge: | Language models struggle to understand and explain the beliefs of others, despite improving performance on a wide variety of tasks. |
| Approach: | They propose to modify the social-chem-101 corpus to allow for perspective-taking, the process of conceptualizing the point of view of another person. |
| Outcome: | The proposed models outperform the recent models conditioned on self-disclosures with high similarity to the conflict situation. |
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models (2025.emnlp-main)
Copied to clipboard
Mehdi Ali, Manuel Brack, Max Lübbering, Elias Wendt, Abbas Goher Khan, Richard Rutmann, Alex Jude, Maurice Kraus, Alexander Arno Weber, Felix Stollenwerk, David Kaczér, Florian Mai, Lucie Flek, Rafet Sifa, Nicolas Flores-Herr, Joachim Koehler, Patrick Schramowski, Michael Fromm, Kristian Kersting
| Challenge: | Existing open-source multilingual datasets rely on heuristic filtering methods restricting both their cross-lingual transferability and scalability. |
| Approach: | They propose a systematic approach that curates diverse and high-quality multilingual data at scale while significantly reducing computational demands. |
| Outcome: | Evaluated empirically across 35 languages, the proposed approach outperforms current heuristic filtering methods like Fineweb2 and improves model training quality and retention rates. |
The Impact of Differential Privacy on Group Disparity Mitigation (2024.findings-naacl)
Copied to clipboard
| Challenge: | a recent study evaluated the impact of differential privacy on fairness across four tasks. |
| Approach: | They evaluate the impact of differential privacy on fairness across four diverse tasks . they train (,)-differentially private models with empirical risk minimization . |
| Outcome: | The proposed model shows that differential privacy increases performance differences between groups . the model also reduces performance differences in the robust setting . |
A Critical Reflection and Forward Perspective on Empathy and Natural Language Processing (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Empathy recognition and empathetic response generation tasks are well-established research directions, but there is little clarity on what empathy is and how it is being operationalized. |
| Approach: | They argue that current directions will benefit from a clear conceptualization that includes operationalizing cognitive empathy components. |
| Outcome: | The proposed framework will help to define and operationalize empathy in natural language processing. |
Suicide Ideation Detection via Social and Temporal User Representations using Hyperbolic Learning (2021.naacl-main)
Copied to clipboard
| Challenge: | Recent studies indicate that individuals exhibiting suicidal ideation increasingly turn to social media rather than mental health practitioners. |
| Approach: | They propose a framework leveraging a user’s emotional history and social information from a users neighborhood in a network to contextualize the interpretation of the latest tweet of a Twitter user. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on this task, showing the benefits of both socially and personally contextualized representations. |
Investigating User Radicalization: A Novel Dataset for Identifying Fine-Grained Temporal Shifts in Opinion (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing models that model fine-grained opinion shifts of social media users are lacking . lack of publicly available datasets for this task presents a major challenge . |
| Approach: | They propose an annotated social media opinion dataset that provides a model for subtle opinion fluctuations and fine-grained stances. |
| Outcome: | The proposed dataset is comparable to the annotations of experts and non-experts. |
DeFaktS: A German Dataset for Fine-Grained Disinformation Detection through Social Media Framing (2024.lrec-main)
Copied to clipboard
| Challenge: | Distinctively curated across various news topics, DeFaktS offers an unparalleled insight into disinformation’s diverse characteristics. |
| Approach: | They propose to annotate every structural component and semantic element of a news piece, eliminating the need for external knowledge sources. |
| Outcome: | The proposed dataset contains 105,855 posts with 20,008 meticulously labeled tweets and eliminates the need for external knowledge sources. |
PHASE: Learning Emotional Phase-aware Representations for Suicide Ideation Detection on Social Media (2021.eacl-main)
Copied to clipboard
| Challenge: | Recent studies indicate that individuals exhibiting suicidal ideation increasingly turn to social media rather than mental health practitioners. |
| Approach: | They propose a time-and-phase-aware framework that adaptively learns features from a user’s historical emotional spectrum to contextualize suicidal intent. |
| Outcome: | The proposed framework outperforms state-of-the-art methods while outperforming existing methods. |
More Agents Improve Math Problem Solving but Adversarial Robustness Gap Persists (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are powerful but weak when inputs are perturbed. |
| Approach: | They evaluate LLMs that are more powerful than single LLM in math question answering . they use a unified sampling-and-voting framework to evaluate their models . |
| Outcome: | The proposed models show that collaboration between agents improves accuracy and clean accuracy even with a large number of agents. |
Perceived and Intended Sarcasm Detection with Graph Attention Networks (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing sarcasm detection systems focus on exploiting linguistic markers, context, or user-level priors, but social studies suggest that the relationship between the author and the audience can be equally relevant for the sarkasmal usage and interpretation. |
| Approach: | They propose a framework leveraging a user context from their historical tweets together with social information from a users neighborhood in an interaction graph to contextualize the interpretation of the post. |
| Outcome: | The proposed framework combines a user context from their historical tweets with social information from a users neighborhood in an interaction graph to contextualize the interpretation of the post. |
On the Limitations of Language-targeted Pruning: Investigating the Calibration Language Impact in Multilingual LLM Pruning (2026.tacl-1)
Copied to clipboard
| Challenge: | Recent advances in large language model pruning have shown high predictive performance in post-training settings. |
| Approach: | They conduct an empirical study on the performance and internal representation changes associated with pruning multilingual models for monolingual applications. |
| Outcome: | The proposed pruning methods retain perplexity and yield high signal-to-noise ratios, but not consistently improve downstream tasks. |
HypMix: Hyperbolic Interpolative Data Augmentation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for data augmentation involve performing mathematical operations over the raw input samples or their latent states representations, but these operations are performed in the Euclidean space, simplifying these representations and resulting in noisy interpolations. |
| Approach: | They propose a model-, data-, and modality-agnostic interpolative data augmentation technique operating in the hyperbolic space that captures the complex geometry of input and hidden state hierarchies better than its contemporaries. |
| Outcome: | The proposed technique outperforms state-of-the-art methods on benchmark and low resource datasets across speech, text, and vision modalities. |
DMix: Adaptive Distance-aware Interpolative Mixup (2022.acl-short)
Copied to clipboard
| Challenge: | Interpolation-based regularisation methods such as Mixup have shown to be effective for various tasks and modalities. |
| Approach: | They propose an adaptive distance-aware interpolative Mixup that selects samples based on their diversity in the embedding space. |
| Outcome: | The proposed method achieves state-of-the-art on sentence classification over existing methods on 8 benchmark datasets across English, Arabic, Turkish, and Hindi languages while achieving benchmark F1 scores in 3 times less number of iterations. |
LeadEmpathy: An Expert Annotated German Dataset of Empathy in Written Leadership Communication (2024.lrec-main)
Copied to clipboard
| Challenge: | Empathetic leadership communication is associated with a wide range of positive individual and organizational outcomes. |
| Approach: | They propose an expert-annotated german dataset for modeling empathy in written leadership communication that uses a theory-based coding scheme to model cognitive and affective empathy in asynchronous communication. |
| Outcome: | The proposed model can be applied to produce high-quality, multidimensional empathy ratings in current and future applications. |
The Practical Impacts of Theoretical Constructs on Empathy Modeling (2025.emnlp-main)
Copied to clipboard
| Challenge: | Empathy operationalizations in NLP are varied, with some having specific behaviors and properties, while others are more abstract. |
| Approach: | They analyze the transfer performance of empathy models adapted to empathy tasks with different theoretical groundings and characterize them as direct, abstract, or adjacent. |
| Outcome: | The proposed models show that they are more transferable than other models. |
USDC: A Dataset of ̲User ̲Stance and ̲Dogmatism in Long ̲Conversations (2025.findings-acl)
Copied to clipboard
| Challenge: | Previously, studies on stance and dogmatism in user conversations have focused on training models using annotated datasets at the post level, treating each post as independent and randomly sampling posts from conversation threads. |
| Approach: | They build a dataset for studying user opinion fluctuations in 764 long multi-user Reddit conversation threads, called USDC. |
| Outcome: | The proposed dataset analyzes user opinion fluctuations in 764 long multi-user Reddit conversation threads. |
Multi-Hop Reasoning for Question Answering with Hyperbolic Representations (2025.findings-acl)
Copied to clipboard
| Challenge: | a rigorous and detailed comparison of the two spaces for multi-hop reasoning is lacking. |
| Approach: | They compare the capacity of hyperbolic space versus Euclidean space in multi-hop reasoning . they use an encoder-decoder model to integrate hyperbolical representations with a knowledge graph . |
| Outcome: | The proposed model outperforms the Euclidean space in multi-hop reasoning. |