Papers by Gabriella Lapesa

25 papers
Moderation in the Wild: Investigating User-Driven Moderation in Online Discussions (2024.eacl-long)

Copied to clipboard

Challenge: Effective content moderation is imperative for fostering healthy and productive discussions in online domains.
Approach: They propose to document and release a dataset of comments in which users act as moderators.
Outcome: The proposed dataset contains 1000 comment-reply pairs with crowdsourced annotations from a large annotator pool and fine-grained annotation schema targeting the functions of moderation, stylistic properties(aggressiveness, subjectivity, sentiment), constructiveness, and individual perspectives of the annotators on the task.
Scaling up Discourse Quality Annotation for Political Science (2022.lrec-1)

Copied to clipboard

Challenge: Existing annotations on deliberative quality are time-consuming and suffer from class imbalance . ephd thesis: deliberation is not only the output of the decision making, but also the discussion that leads up to it.
Approach: They propose to use data augmentation techniques to improve deliberative quality predictions in a standard dataset.
Outcome: The proposed methods outperform classifiers based on linguistic features and argument quality annotations with or without data augmentation.
DEbateNet-mig15:Tracing the 2015 Immigration Debate in Germany Over Time (2020.lrec-1)

Copied to clipboard

Challenge: a dataset for germany covering the public debate on immigration is annotated . a political science notion of a claim is used to represent the political discourse .
Approach: They annotate a dataset for german public debate on immigration in 2015 using a political science notion of a claim . they identify claims in newspaper articles, assign them to actors and fine-grained categories and annotize their polarity and date.
Outcome: The dataset is annotated by a political science framework and shows it captures political debate . it shows that political actors can change their positions and take a strong stand against them .
Argument Quality Assessment in the Age of Instruction-Following Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Argument quality assessment is critical for opinion formation, decision making, writing education, and the like.
Approach: They propose to use large language models to leverage knowledge across contexts to enable a much more reliable assessment.
Outcome: The proposed approach improves the quality of argumentation and the ability to leverage knowledge across contexts.
Node Placement in Argument Maps: Modeling Unidirectional Relations in High & Low-Resource Scenarios (2023.acl-long)

Copied to clipboard

Challenge: Argument maps structure discourse into nodes with each node being an argument that supports or opposes its parent argument.
Approach: They propose a task of node placement: suggesting candidate nodes as parents for a new contribution.
Outcome: The proposed method improves the quality of the argument maps and reduces redundancy.
Towards a Perspectivist Turn in Argument Quality Assessment (2025.naacl-long)

Copied to clipboard

Challenge: Argument quality is a key aspect of computational argumentation (CA), but it still exhibits a high degree of subjectivity in perception.
Approach: They propose to use a multi-layered classification to target two aspects of argument quality in a systematic review of NLP datasets.
Outcome: The proposed model improves the quality of annotators and their ability to be used in perspectivist research.
StoryARG: a corpus of narratives and personal experiences in argumentative texts (2023.acl-long)

Copied to clipboard

Challenge: Narratives and argumentation are deeply related, according to psychologists and social scientists.
Approach: They annotated StoryARG from well-established corpora in computational argumentation and the Social Sciences, as well as comments to New York Times articles.
Outcome: The dataset contains 2451 textual spans annotated at two levels . it reveals positive impact on effectiveness for stories which illustrate a solution to a problem and in general, annotator-specific preferences .
Bridging Argument Quality and Deliberative Quality Annotations with Adapters (2023.findings-eacl)

Copied to clipboard

Challenge: Assessing the quality of an argument is a complex, highly subjective task . argument quality dimensions are complex and dependent on the context in which it is assessed .
Approach: They propose a multi-task learning framework that incorporates knowledge about related dimensions into the learning process.
Outcome: The proposed framework improves quality prediction in an extrinsic, out-of-domain task.
Tell Me What You Know About Sexism: Expert-LLM Interaction Strategies and Co-Created Definitions for Zero-Shot Sexism Detection (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) with chat interfaces are increasingly popular in various scientific fields, for a variety of tasks related to social science research questions.
Approach: They propose to use large language models to combine human and machine expertise to improve their models' performance.
Outcome: The proposed model performs better with co-created definitions than with expert-written definitions.
It Is Not Only the Negative that Deserves Attention! Understanding, Generation & Evaluation of (Positive) Moderation (2025.naacl-long)

Copied to clipboard

Challenge: Moderation is essential for maintaining and improving the quality of online discussions.
Approach: They annotate a dataset on 13 modes of discussion and use it to generate positive moderation.
Outcome: The proposed model shows that professional moderation generates higher ratings than professional moderated moderation, but prefers professional moderate in pairwise comparison.
Improving Neural Political Statement Classification with Class Hierarchical Information (2022.findings-acl)

Copied to clipboard

Challenge: skewed classification of fine-grained categories in text-based computational social science is challenging on the NLP side.
Approach: They propose to use hierarchical relations among categories in codebooks to create constraints on the learned model.
Outcome: The proposed model improves on two datasets and multiple languages.
Mining the uncertainty patterns of humans and models in the annotation of moral foundations and human values (2025.acl-long)

Copied to clipboard

Challenge: disagreement in annotation (HLV) is considered a constitutive feature of subjective tasks.
Approach: They investigate the relationship between disagreement in annotation and model uncertainty . they use linguistic features to calibrate models to HLV and uncertainty to analyze their impact on uncertainty.
Outcome: The proposed model uncertainty is calibrated to human label variation (HLV) the proposed model is calibrate to human labels, the authors show .
Self-reported Demographics and Discourse Dynamics in a Persuasive Online Forum (2024.lrec-main)

Copied to clipboard

Challenge: Research on language as interactive discourse demonstrates the deliberate use of demographic parameters such as gender, ethnicity, and class to shape social identities.
Approach: They propose to investigate the role and effects of gender self-disclosures on online discourse dynamics by focusing on author gender.
Outcome: The proposed dataset will provide a further impulse for research on the interplay between gender disclosures, community interaction, and persuasion in online discourse.
Investigating Independence vs. Control: Agenda-Setting in Russian News Coverage on Social Media (2022.lrec-1)

Copied to clipboard

Challenge: a major challenge in the media industry has always been its targeted manipulation, says a new study . agenda-setting is a well-known phenomenon in political science . authors explore the relationship between economic indicators and mentions of foreign geopolitical entities .
Approach: They investigate agenda-setting in the Russian social media landscape . they explore the relation between economic indicators and mentions of foreign geopolitical entities .
Outcome: The authors examine the relationship between economic indicators and mentions of foreign geopolitical entities, as well as of Russia itself.
Mining, Assessing, and Improving Arguments in NLP and the Social Sciences (2024.lrec-tutorials)

Copied to clipboard

Challenge: a tutorial on computational argumentation is updated to address the problem of argument quality . argument quality is a field of interdisciplinary research that connects natural language processing to social sciences .
Approach: They present an updated version of the EACL 2023 tutorial on argument quality . they will focus on the notions of argument quality across disciplines .
Outcome: The updated version of the EACL 2023 tutorial focuses on argument quality assessment . the authors will focus on the interface between Argument Mining and Deliberation Theory .
Mining, Assessing, and Improving Arguments in NLP and the Social Sciences (2023.eacl-tutorials)

Copied to clipboard

Challenge: a tutorial on argument quality assessment will focus on what makes an argument good or bad . argument quality is a field encompassing varying tasks on the automated analysis and synthesis of natural language arguments.
Approach: This tutorial will focus on the assessment of argument quality across disciplines . authors will involve participants in annotation studies on the quality assessment .
Outcome: The tutorial will focus on the assessment of argument quality across disciplines . it will involve participants in two annotation studies on the quality assessment and the improvement of quality .
From Emotion to Expression: Theoretical Foundations and Resources for Fear Speech (2026.eacl-long)

Copied to clipboard

Challenge: a new study of fear speech is under-resourced and fragmented. authors review existing definitions and propose a taxonomy that consolidates different dimensions of fear.
Approach: They propose a taxonomy that consolidates different dimensions of fear for studying fear speech.
Outcome: The proposed taxonomy consolidates different dimensions of fear for studying fear speech.
An Environment for Relational Annotation of Political Debates (P19-3)

Copied to clipboard

Challenge: Scalable text analysis techniques can open corpora to new questions in computational social sciences and digital humanities.
Approach: They describe a tool that allows annotating newspaper text with rich information about claims (demands) raised by politicians and other actors.
Outcome: The MARDY tool realizes the complete workflow necessary for annotating a large newspaper text collection with rich information about claims (demands) raised by politicians and other actors.
Towards Argument Mining for Social Good: A Survey (2021.acl-long)

Copied to clipboard

Challenge: Argument Mining is a social science-based approach to analysis and analysis of arguments.
Approach: They propose a novel definition of argument quality which integrates the social science literature and the argument quality.
Outcome: The proposed definition of argument quality integrates the social science literature and the argument quality debate.
Reports of personal experiences and stories in argumentation: datasets and analysis (2022.acl-long)

Copied to clipboard

Challenge: Personal experiences and stories are important in argumentation, but they are not considered in the social sciences.
Approach: They propose to use annotated documents to scale-up the analysis using existing annotations.
Outcome: The proposed classifiers can identify documents containing personal experiences and reports . they can scale up to three domains and show that they perform well across domains.
Toeing the Party Line: Election Manifestos as a Key to Understand Political Discourse on Twitter (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent work on political positioning on Twitter has tended to focus on manifestos rather than social media since it is ambiguous and dependent on social context.
Approach: They propose to use hashtags as a signal to fine-tune text representations for politicians' tweets using a hashtag-based method to predict pairwise positional similarities between parties from the manifesto case to the Twitter case.
Outcome: The proposed method matches politicians' statements to official lines of the parties' tweets, even when only small subsets from shorter time periods are available.
Stories and Personal Experiences in the COVID-19 Discourse (2024.lrec-main)

Copied to clipboard

Challenge: 'storytelling' is a human strategy to use personal experiences to back-up one's position in debates about controversial topics.
Approach: They analyse the use of storytelling in the COVID-19 discourse by automatically annotating three publicly available Reddit datasets for a total of 367K comments.
Outcome: The proposed analysis on three publicly available Reddit datasets shows that storytelling is a powerful argumentative tool.
AI Argues Differently: Distinct Argumentative and Linguistic Patterns of LLMs in Persuasive Contexts (2025.emnlp-main)

Copied to clipboard

Challenge: Distinguishing LLM-generated text from human-written is a key challenge for safe and ethical NLP, especially in high-stake settings such as persuasive online discourse.
Approach: They propose to use general-purpose linguistic features and domain-specific features related to argument quality to compare human- and LLM-authored arguments.
Outcome: The proposed framework compares arguments by humans and three LLMs using two easily-interpretable feature sets.
PerspectiveMod: A Perspectivist Resource for Deliberative Moderation (2025.emnlp-main)

Copied to clipboard

Challenge: Human moderators in online discussions face a heterogeneous range of tasks that go beyond content moderation, or policing.
Approach: They propose a dataset of online comments annotated for the question "Does this comment require moderation?" they aim to improve discussion quality by analyzing annotator perspectives and annotating their views.
Outcome: The proposed model is unique in its intentional variation across the level of moderation experience embedded in the source data, the annotator profiles and the individuality of the annnotator.
How to Translate Your Samples and Choose Your Shots? Analyzing Translate-train & Few-shot Cross-lingual Transfer (2022.findings-naacl)

Copied to clipboard

Challenge: Recent studies have focused on zero-shot cross-lingual transfer of pretrained languages.
Approach: They propose to use few-shot cross-lingual transfer to improve zero-shot performance of multilingual pretrained language models.
Outcome: The proposed model can be scaled to high-quality samples and improves on zero-shot performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations