Papers by Serena Villata
Argument-based Detection and Classification of Fallacies in Political Debates (2023.emnlp-main)
Copied to clipboard
| Challenge: | Fallacies are arguments that employ faulty reasoning, causing inaccurate conclusions and invalid inferences . ad hominem fallacy is one of the most common fallacy labels used in political debates despite its use in many scenarios . |
| Approach: | They extend the ElecDeb60To16 dataset of U.S. presidential debates annotated with fallacious arguments by incorporating the most recent Trump-Biden debate. |
| Outcome: | The proposed method extends the ElecDeb60To16 dataset of U.S. presidential debates annotated with fallacious arguments . |
MedMT5: An Open-Source Multilingual Text-to-Text LLM for the Medical Domain (2024.lrec-main)
Copied to clipboard
Iker García-Ferrero, Rodrigo Agerri, Aitziber Atutxa Salazar, Elena Cabrio, Iker de la Iglesia, Alberto Lavelli, Bernardo Magnini, Benjamin Molinet, Johana Ramirez-Romero, German Rigau, Jose Maria Villa-Gonzalez, Serena Villata, Andrea Zaninello
| Challenge: | Existing studies on large language models for medical applications have focused on a single language . medical mT5 outperforms both encoders and similar sized text-to-text models in English, French, and Italian benchmarks . |
| Approach: | They propose to train Medical mT5, the first open-source text-to-text multilingual model for the medical domain. |
| Outcome: | The proposed model outperforms encoders and similar sized models on the Spanish, French, and Italian benchmarks while being competitive with current state-of-the-art models in English. |
AM4DSP: Argumentation Mining in Structured Decentralized Discussion Platforms for Deliberative Democracy (2025.emnlp-demos)
Copied to clipboard
| Challenge: | Argument mining is the automated process of identification and extraction of argumentative structures in natural language. |
| Approach: | They propose to use argument mining to extract arguments from online discussions in the context of deliberative democracy. |
| Outcome: | The proposed system enables the extraction and analysis of arguments from online discussions in the context of deliberative democracy. |
Argument Quality Assessment in the Age of Instruction-Following Large Language Models (2024.lrec-main)
Copied to clipboard
Henning Wachsmuth, Gabriella Lapesa, Elena Cabrio, Anne Lauscher, Joonsuk Park, Eva Maria Vecchi, Serena Villata, Timon Ziegenbein
| Challenge: | Argument quality assessment is critical for opinion formation, decision making, writing education, and the like. |
| Approach: | They propose to use large language models to leverage knowledge across contexts to enable a much more reliable assessment. |
| Outcome: | The proposed approach improves the quality of argumentation and the ability to leverage knowledge across contexts. |
Hybrid Emoji-Based Masked Language Models for Zero-Shot Abusive Language Detection (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies have demonstrated the effectiveness of cross-lingual language model pre-training on NLP tasks. |
| Approach: | They propose a hybrid emoji-based Masked Language Model to leverage eojis across languages to improve the learning of short text messages. |
| Outcome: | The proposed model performs better on German, Italian and Spanish. |
Graph Embeddings for Argumentation Quality Assessment (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Argumentation is the process by which arguments are constructed, compared, evaluated in several respects and judged in order to establish whether any of them is warranted. |
| Approach: | They propose to annotate 1908 arguments tagged with quality facets from a resource of 402 persuasive essays and to use them to create a neural architecture that takes into account the support and attack relations holding among the arguments. |
| Outcome: | The proposed neural architecture outperforms state-of-the-art and standard arguments on the persuasive essays dataset. |
An In-depth Analysis of Implicit and Subtle Hate Speech Messages (2023.eacl-main)
Copied to clipboard
| Challenge: | Explicit hate speech is more easily identifiable by recognizing hateful words, but subtle messages are harmful . subtle messages contain linguistically subtle and implicit forms of HS, such as circumlocution, metaphors and sarcasm . social media have faced pressure from civil rights groups demanding to monitor and limit online hate speech . |
| Approach: | They propose to use a fine-grained definition of implicit and subtle messages to detect HS . they then experiment with neural network architectures to detect subtle content . |
| Outcome: | The proposed models perform satisfactory on explicit messages, but fail to detect subtle content. |
Increasing Argument Annotation Reproducibility by Using Inter-annotator Agreement to Improve Guidelines (L18-1)
Copied to clipboard
| Challenge: | Argument Mining systems require large amounts of data to characterize phenomena and find patterns that can be exploited by an automatic analyzer. |
| Approach: | They propose to exploit inter-annotator agreement measures to improve Argument annotation guidelines. |
| Outcome: | The proposed method improves Argument annotation guidelines by exploiting inter-annotator agreement measures. |
DISPUTool 3.0: Fallacy Detection and Repairing in Argumentative Political Debates (2025.acl-demo)
Copied to clipboard
| Challenge: | DISPUTool 3.0 is a web-based application for identifying and fixing fallacious arguments in political debates. |
| Approach: | They propose a web-based application designed to identify and repair fallacious arguments in political debates. |
| Outcome: | The proposed tool is based on the ElecDeb60to20 dataset covering US presidential debates from 1960 to 2020. |
CyberAgressionAdo-v1: a Dataset of Annotated Online Aggressions in French Collected through a Role-playing Game (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent studies have highlighted that private instant messaging platforms are major mediums of cyber aggression among teens. |
| Approach: | They present a dataset of aggressive chats in French collected through a role-playing game in high-schools . they provide insights on the different types of aggression and verbal abuse depending on the targeted victims . |
| Outcome: | The proposed dataset analyzes aggressive conversations in French on a role-playing game in high schools . it provides insights on the different types of aggression and verbal abuse depending on the targeted victims . |
Mining, Assessing, and Improving Arguments in NLP and the Social Sciences (2024.lrec-tutorials)
Copied to clipboard
| Challenge: | a tutorial on computational argumentation is updated to address the problem of argument quality . argument quality is a field of interdisciplinary research that connects natural language processing to social sciences . |
| Approach: | They present an updated version of the EACL 2023 tutorial on argument quality . they will focus on the notions of argument quality across disciplines . |
| Outcome: | The updated version of the EACL 2023 tutorial focuses on argument quality assessment . the authors will focus on the interface between Argument Mining and Deliberation Theory . |
Mining, Assessing, and Improving Arguments in NLP and the Social Sciences (2023.eacl-tutorials)
Copied to clipboard
| Challenge: | a tutorial on argument quality assessment will focus on what makes an argument good or bad . argument quality is a field encompassing varying tasks on the automated analysis and synthesis of natural language arguments. |
| Approach: | This tutorial will focus on the assessment of argument quality across disciplines . authors will involve participants in annotation studies on the quality assessment . |
| Outcome: | The tutorial will focus on the assessment of argument quality across disciplines . it will involve participants in two annotation studies on the quality assessment and the improvement of quality . |
Regrexit or not Regrexit: Aspect-based Sentiment Analysis in Polarized Contexts (2020.coling-main)
Copied to clipboard
| Challenge: | Aspect-based Sentiment Analysis (ABSA) aims at capturing sentiment expressed toward each aspect of a target entity. |
| Approach: | They propose to extend the task of Aspect-based Sentiment Analysis (ABSA) toward affect and emotion representation in polarized settings. |
| Outcome: | The proposed model captures aspect-based polarization from newspapers regarding the Brexit scenario of 1.2m entities at sentence-level. |
Playing the Part of the Sharp Bully: Generating Adversarial Examples for Implicit Hate Speech Detection (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing algorithms for hate speech detection focus on explicit forms of hate speech, but they fail to properly detect subtle and implicit HS messages. |
| Approach: | They propose a framework for generating adversarial implicit HS short-text messages using Auto-regressive language models and a strategy to group the generated messages in complexity levels. |
| Outcome: | The proposed framework shows that iteratively retraining on HARD messages significantly improves implicit HS benchmarks. |
Yes, we can! Mining Arguments in 50 Years of US Presidential Campaign Debates (P19-1)
Copied to clipboard
| Challenge: | Political debates are a natural application scenario for Argument Mining. |
| Approach: | They propose an argument mining approach to political debates that uses argument components to annotate 39 political debate from the last 50 years of US presidential campaigns. |
| Outcome: | The proposed approach outperforms baselines in argument mining over political debates. |
CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative Structures (2024.emnlp-main)
Copied to clipboard
Ekaterina Sviridova, Anar Yeginbergen, Ainara Estarrona, Elena Cabrio, Serena Villata, Rodrigo Agerri
| Challenge: | Existing tools to aid residents in teaching medical doctors to explain decisions are a key objective of AI in education. |
| Approach: | They present a multilingual dataset for Medical Question Answering where doctors can annotate correct and incorrect diagnoses with argument components and argument relations. |
| Outcome: | The proposed dataset consists of 558 clinical cases with explanations in English, Spanish, French, Italian and annotated with argument components and argument relations. |
Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering (2024.emnlp-main)
Copied to clipboard
| Challenge: | Automated responses lack argumentative richness which characterises expert-produced counterspeech. |
| Approach: | They propose to automate counterspeech generation by investigating tension between helpfulness and harmlessness of LLMs and to assess whether presence of safety guardrails hinders quality of generations. |
| Outcome: | The proposed approach produces more cogent responses that lack argumentative richness which characterises expert-produced counterspeech. |
Unmasking the Hidden Meaning: Bridging Implicit and Explicit Hate Speech Embedding Representations (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to detect explicit hate speech (HS) are focusing on detecting explicit forms of hateful expressions on user-generated content. |
| Approach: | They propose to examine the differences between embedding implicit and explicit hateful messages . they compare and link explicit and implicit hateful message across datasets . |
| Outcome: | The proposed model improves on explicit hate speech detection while retaining high performance on borderline cases. |