Papers by Lucie-Aimée Kaffee

7 papers
Investigating Human Values in Online Communities (2025.naacl-long)

Copied to clipboard

Challenge: Existing value frameworks struggle with sample sizes and rely on selfreported surveys to calculate values.
Approach: They propose a method to computationally analyse values on Reddit using in-domain and out-of-domain human annotations to train a value relevance and a polarity classifier.
Outcome: The proposed method can be used to analyse values on reddit using human annotations and human annotation.
Why Should This Article Be Deleted? Transparent Stance Detection in Multilingual Wikipedia Editor Discussions (2023.emnlp-main)

Copied to clipboard

Challenge: Currently, only 20% of the English comments explicitly mention content moderation policies, but as few as 2% of the German and Turkish comments.
Approach: They propose to use a multilingual dataset to predict stances with existing content moderation policies and to use them to explain moderation decisions.
Outcome: The proposed model predicts stances and corresponding reasons with high accuracy, adding transparency to the decision-making process.
Learning to Generate Wikipedia Summaries for Underserved Languages from Wikidata (N18-2)

Copied to clipboard

Challenge: Existing Wikipedia content is unevenly distributed among 287 languages . authors propose a neural network architecture that generates textual summaries from Wikidata triples .
Approach: They propose an automated approach to generate Wikipedia summaries from Wikidata triples using structured data.
Outcome: The proposed approach is tested on Arabic and Esperanto languages with limited editors and content in the most under-resourced Wikipedias.
Investigating Wit, Creativity, and Detectability of Large Language Models in Domain-Specific Writing Style Adaptation of Reddit’s Showerthoughts (2024.starsem-1)

Copied to clipboard

Challenge: Recent Large Language Models (LLMs) have shown the ability to generate content that is difficult or impossible to distinguish from human writing.
Approach: They compare GPT-2 and GPT-Neo fine-tuned on Reddit data and GTP-3.5 invoked in a zero-shot manner, against human-authored texts.
Outcome: The proposed model can generate short, creative texts that are difficult to distinguish from human writing, but human evaluators rate them worse than the model.
Thorny Roses: Investigating the Dual Use Dilemma in Natural Language Processing (2023.findings-emnlp)

Copied to clipboard

Challenge: Dual use is a problem in the context of natural language processing, says aaron eliotta . eelisa et al.: it is important to examine their rightful use and potential misuse .
Approach: They propose a definition and checklist for dual-use in natural language processing based on a survey of NLP researchers and practitioners.
Outcome: The proposed checklist focuses on dual-use in NLP based on a survey of NLP researchers and practitioners.
Presumed Cultural Identity: How Names Shape LLM Responses (2025.findings-emnlp)

Copied to clipboard

Challenge: Names can be used as markers of individuality, cultural heritage, and personal history when interacting with chatbots.
Approach: They propose to use names as cultural bias in chatbots to adapt to user input and task contexts.
Outcome: The proposed method demonstrates that LLMs make cultural identity assumptions based on their users’ presumed backgrounds based upon their names .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations