Papers by Alberto Barrón-Cedeño

17 papers
LLMs Underperform Graph-Based Parsers on Supervised Relation Extraction for Complex Graphs (2026.acl-short)

Copied to clipboard

Challenge: Relation extraction is a core NLP task which involves extracting [head, relation, dependent] RDF triples from text.
Approach: They evaluate four large language models against a graph-based parser on six relation extraction datasets with sentence graphs of varying sizes and complexities.
Outcome: The graph-based parser outperforms the LLMs on six relation extraction datasets with sentence graphs of varying sizes and complexities.
Elote, Choclo and Mazorca: on the Varieties of Spanish (2024.naacl-long)

Copied to clipboard

Challenge: Spanish is the official language in 20 countries and the second most-spoken native language . available corpora treat it as one monolithic language, damping prediction power .
Approach: They compile and curate datasets in different varieties of Spanish around the world at an unprecedented scale and create the CEREAL corpus.
Outcome: The results show that Spanish is a multilingual language with a wide range of cultural and cultural influences.
Untangling Hate Speech Definitions: A Semantic Componential Analysis Across Cultures and Domains (2025.findings-naacl)

Copied to clipboard

Challenge: a new framework for analyzing hate speech definitions is proposed to address cultural differences in interpretations . a dataset of 493 definitions from more than 100 cultures is used to analyze hate speech .
Approach: They propose a framework for a cross-cultural and cross-domain analysis of hate speech definitions . they use open-source LLMs to analyze the impact of different definitions on hate speech detection .
Outcome: The proposed framework enables cross-cultural and cross-domain analysis of hate speech definitions . it reveals that many domains borrow definitions from one another without taking into account target culture .
A Corpus for Sentence-Level Subjectivity Detection on English News Articles (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to spotting subjectivity require language-specific tools.
Approach: They develop annotation guidelines for sentence-level subjectivity detection that are not limited to language-specific cues.
Outcome: The proposed framework enables subjectivity detection in English and across other languages without relying on language-specific tools, such as lexicons or machine translation.
The (Undesired) Attenuation of Human Biases by Multilinguality (2022.emnlp-main)

Copied to clipboard

Challenge: odor pleasantness perception is universal, but cultural biases are not always present in embedding models . et al., 2018: a new study shows that cultural bias is not always the case in embedded models based on human texts .
Approach: They propose multilingual cultural aware tests to quantify biases in embedding models . they find that biased models are more likely to be multilingual than monolingual ones .
Outcome: The results show that human preferences are not always universal . they also show that multilinguality reverses biases, despite differences in training corpus .
Misogyny and Aggressiveness Tend to Come Together and Together We Address Them (2022.lrec-1)

Copied to clipboard

Challenge: Using a binary task to identify whether a tweet is misogynous and aggressive, we compare two approaches to address these problems: one multi-class model that discriminates between all the classes at once; and a cascaded approach where the binary classification is carried out separately.
Approach: They propose a multi-class model that discriminates between all the classes at once and a cascaded approach where the binary classification is carried out separately and then joined together.
Outcome: The proposed models outperform the top submissions to Evalita on the 2020 shared task on automatic misogyny and aggressiveness identification in Italian tweets.
The Challenges of Creating a Parallel Multilingual Hate Speech Corpus: An Exploration (2024.lrec-main)

Copied to clipboard

Challenge: Hate speech is one of the most demanding topics in Natural Language Processing, as its multifaceted nature is accompanied by a handful of challenges, such as multilinguality and cross-linguality.
Approach: They propose a pipeline that could be used to create a parallel multilingual hate speech dataset using machine translation.
Outcome: The proposed pipeline will be able to create a parallel multilingual hate speech dataset using machine translation.
Language is Scary when Over-Analyzed: Unpacking Implied Misogynistic Reasoning with Argumentation Theory-Driven Prompts (2024.emnlp-main)

Copied to clipboard

Challenge: a new study aims to understand the implicit reasoning used to convey misogynistic comments in Italian and English.
Approach: They propose misogyny detection as an Argumentative Reasoning task and use argumentation theory to build large language models to understand the implicit reasoning used to convey misogany in Italian and English.
Outcome: The proposed task is an argumentative reasoning task in Italian and English.
Fine-Grained Analysis of Propaganda in News Article (D19-1)

Copied to clipboard

Challenge: Existing methods for detecting propaganda are noisy and lack of explainability.
Approach: They propose to perform fine-grained analysis of texts by detecting all fragments that contain propaganda techniques as well as their type.
Outcome: The proposed model outperforms several strong BERT-based baselines.
A Checkpoint on Multilingual Misogyny Identification (2022.acl-srw)

Copied to clipboard

Challenge: a study on hate speech against minorities in Italian tweets found that 1 women are the most targeted group.
Approach: They propose to train monolingual transformers and multilingual transformer models with monolingual data in English, Italian, and Spanish to detect misogyny in tweets.
Outcome: The proposed model achieves state-of-the-art on English, Italian, and Spanish.
Prta: A System to Support the Analysis of Propaganda Techniques in the News (2020.acl-demos)

Copied to clipboard

Challenge: recent events have brought the public attention to the dangers of online disinformation.
Approach: a new tool helps users analyze propaganda using specific rhetorical and psychological techniques. a prta system identifies the spans in which propaganda techniques occur and compares them.
Outcome: a new tool can analyze articles crawled on a regular basis and compare them on the basis of their use of propaganda techniques.
Tanbih: Get To Know What You Are Reading (D19-3)

Copied to clipboard

Challenge: Nowadays, more and more readers consume news online.
Approach: They propose a news platform that displays news grouped into events and generates media profiles that show the general factuality of reporting, the degree of propagandistic content, hyper-partisanship, leading political ideology, general frame of reporting and stance with respect to various claims and topics of a media outlet.
Outcome: The proposed news platform displays news grouped into events and generates media profiles that show the factuality of reporting, the degree of propagandistic content, hyper-partisanship, leading political ideology, general frame of reporting and stance with respect to various claims and topics of a news outlet.
A Flexible, Efficient and Accurate Framework for Community Question Answering Pipelines (P18-4)

Copied to clipboard

Challenge: Community Question Answering is a research area that benefits from deep linguistic analysis . previous cQA challenges have shown that neural approaches are not enough to deliver state-of-the-art results .
Approach: They propose a framework to distribute computation of cQA tasks over computer clusters . community question answering is a research area that benefits from deep linguistic analysis .
Outcome: The proposed framework scales to large datasets and delivers fast processing.
The “r” in “woman” stands for rights. Auditing LLMs in Uncovering Social Dynamics in Implicit Misogyny (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study examined misogynistic expressions in English and Italian . a taxonomy of social dynamics is used to identify misogorical expressions .
Approach: They examine misogynistic expressions in English and Italian using a taxonomy of social dynamics . they find that LLMs struggle to follow instructions and reason in all settings .
Outcome: The results show that misogynistic expressions are more often implicit than openly hostile . the authors show that LLMs struggle to follow instructions and reason in all settings .
Findings of the NLP4IF-2019 Shared Task on Fine-Grained Propaganda Detection (D19-50)

Copied to clipboard

Challenge: A shared task on fine-grained propaganda detection was organized at EMNLP-IJCNLP 2019 . 12 systems submitted systems for the FLC task, 25 for the SLC task, and 14 teams submitted a system description paper .
Approach: They present a task on fine-grained propaganda detection as part of the NLP4IF workshop at EMNLP-IJCNLP 2019 . they used a corpus of news articles annotated with an inventory of propagandist techniques at the fragment level to determine the propaganda technique used in each fragment .
Outcome: The shared task on fine-grained propaganda detection was organized at the EMNLP-IJCNLP 2019 . 12 systems submitted for the FLC task, 25 for the SLC task, and 14 submitted a system description paper .
PejorativITy: Disambiguating Pejorative Epithets to Improve Misogyny Detection in Italian Tweets (2024.lrec-main)

Copied to clipboard

Challenge: Disambiguating the meaning of pejorative words might help misogyny detection . state-of-the-art models struggle to correctly classify misogoyne when sentences contain such terms.
Approach: They present a corpus of 1,200 manually annotated Italian tweets for pejorative language at the word level and misogyny at the sentence level.
Outcome: The proposed model improves on 1,200 manually annotated Italian tweets and on two benchmarks.
ClaimRank: Detecting Check-Worthy Claims in Arabic and English (N18-5)

Copied to clipboard

Challenge: ClaimRank is an online system for detecting check-worthy claims . it can be used to prioritize the claims fact-checkers should consider first .
Approach: ClaimRank is an online system for detecting check-worthy claims . it is originally trained on political debates, but can work for any kind of text . authors propose to make automated fact-checking easier by prioritizing claims based on annotations from reputable fact- checking organizations.
Outcome: ClaimRank is an online system for detecting check-worthy claims . it can mimic the sentence selection strategies of reputable fact-checking organizations .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations