Papers by Valerio Basile
WebNLG-IT: Construction of an aligned RDF-Italian corpus through Machine Translation techniques (2025.findings-acl)
Copied to clipboard
| Challenge: | Using NMT and hand-written rules, we created the first aligned Italian RDF-to-text corpus . |
| Approach: | They propose to use NMT to create an Italian version of the WebNLG corpus and to refine and improve the quality of the produced resource. |
| Outcome: | The proposed system is the best on the original English version and the best in the second step, it improves and refines the quality of the produced resource. |
The DipInfoUniTo Realizer at SRST’19: Learning to Rank and Deep Morphology Prediction for Multilingual Surface Realization (D19-63)
Copied to clipboard
| Challenge: | SR is one of the main tasks involved in Natural Language Generation. |
| Approach: | They propose a system which divides the SR task into two independent subtasks, namely word order prediction and morphology inflection prediction. |
| Outcome: | The proposed system is a direct successor to the architecture presented at SR'19. |
Italian NLP for Everyone: Resources and Models from EVALITA to the European Language Grid (2022.lrec-1)
Copied to clipboard
| Challenge: | European Language Grid enables researchers and practitioners to easily distribute and use NLP resources and models. |
| Approach: | They propose to integrate Italian NLP resources into the European Language Grid . they show how easy it is to use the integrated systems and demonstrate how seamless it is . |
| Outcome: | The European Language Grid enables researchers and practitioners to easily distribute and use NLP resources and models. |
PERSEVAL: A Framework for Perspectivist Classification Evaluation (2025.emnlp-main)
Copied to clipboard
Soda Marem Lo, Silvia Casola, Erhan Sezerer, Valerio Basile, Franco Sansonetti, Antonio Uva, Davide Bernardi
| Challenge: | Perspectivist evaluation practices in NLP remain fragmented and inconsistent . |
| Approach: | They propose a framework that evaluates perspectivist models at the individual annotator level and treats annotators and users as distinct entities, consistent with real-world scenarios. |
| Outcome: | The proposed framework evaluates annotators and users as distinct entities consistent with real-world scenarios. |
Capturing Perspectives of Crowdsourced Annotators in Subjective Learning Tasks (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing approaches to label aggregation fail to capture subjective annotations and can lead to biases. |
| Approach: | They propose annotator-aware representations for text for subjective classification tasks that involve learning representations of annotators. |
| Outcome: | The proposed model improves on metrics that assess the performance on capturing individual annotators’ perspectives. |
Label Augmentation for Zero-Shot Hierarchical Text Classification (2024.acl-long)
Copied to clipboard
| Challenge: | Hierarchical Text Classification is a difficult problem due to the lack of labeled data and the cost of manually annotating data samples. |
| Approach: | They propose a method that uses a Large Language Model to augment the deepest layer of the labels hierarchy to enhance its specificity. |
| Outcome: | The proposed method achieves state-of-the-art on four public datasets and a strong correlation between the metric values and the classification performance. |
QUEEREOTYPES: A Multi-Source Italian Corpus of Stereotypes towards LGBTQIA+ Community Members (2024.lrec-main)
Copied to clipboard
Alessandra Teresa Cignarella, Manuela Sanguinetti, Simona Frenda, Andrea Marra, Cristina Bosco, Valerio Basile
| Challenge: | a dataset of social media texts addressing LGBTQIA+ individuals is presented in this paper . the dataset is based on two sources in italian: Facebook and Twitter . |
| Approach: | They describe a dataset composed of two sub-corpora from two different sources in Italian . the dataset includes social media texts regarding LGBTQIA+ individuals, behaviors, ideology and events . |
| Outcome: | The QUEEREOTYPES dataset includes social media texts regarding LGBTQIA+ individuals, behaviors, ideology and events. |
I’m sure you’re a real scholar yourself: Exploring Ironic Content Generation by Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Moreover, irony is highly subjective and can depend on various factors, such as social, cultural, or generational aspects. |
| Approach: | They propose to fine-tune two large language models to generate ironic and non-ironic content and analyze their outputs from a linguistic perspective. |
| Outcome: | The proposed models generate ironic and non-ironic responses to a given social media post and analyze their outputs from a linguistic perspective. |
EPIC: Multi-Perspective Annotation of a Corpus of Irony (2023.acl-long)
Copied to clipboard
Simona Frenda, Alessandro Pedrani, Valerio Basile, Soda Marem Lo, Alessandra Teresa Cignarella, Raffaella Panizzon, Cristina Marco, Bianca Scarlini, Viviana Patti, Cristina Bosco, Davide Bernardi
| Challenge: | EPIC is the first annotated corpus for irony analysis based on data perspectivism . a recent trend in natural language processing (NLP) postulates that the disagreement among annotators in a language resource is a valuable source of knowledge, rather than noise that ought to be minimized or discarded. |
| Approach: | They propose to annotate an English perspectivist irony corpus based on data perspectivism . they validate the model by creating perspective-aware models that encode the perspectives of annotators grouped according to their demographic characteristics. |
| Outcome: | The proposed model can capture different perspectives on irony among different groups of annotators, and is more confident than non-perspectivist models. |
Confidence-based Ensembling of Perspective-aware Models (2023.emnlp-main)
Copied to clipboard
Silvia Casola, Soda Lo, Valerio Basile, Simona Frenda, Alessandra Cignarella, Viviana Patti, Cristina Bosco
| Challenge: | Human label variability has been a topic of research in the field of NLP recently . Exploiting disagreements in annotations has been shown to offer advantages for accurate modelling and fairer evaluation. |
| Approach: | They propose a highly perspectivist model that exploits disagreements in annotations to capture the subjectivity encoded in the annotation process. |
| Outcome: | The proposed model is validated on irony and hate speech detection scenarios in in-domain and cross-domain settings. |
UINAUIL: A Unified Benchmark for Italian Natural Language Understanding (2023.acl-demo)
Copied to clipboard
| Challenge: | a benchmark of six tasks for Italian Natural Language Understanding is presented . large language models (LLMs) have revolutionized the field of natural language processing . a few benchmarks exist for non-English languages, but only a handful are available for nonEnglish languages . |
| Approach: | They introduce a benchmark for Italian Natural Language Understanding that harmonizes the data format and exposes functionalities to facilitate data manipulation and evaluation of custom models. |
| Outcome: | The proposed benchmarks are based on the European Language Grid and available models in Italian and multilingual languages. |
Do You Really Want to Hurt Me? Predicting Abusive Swearing in Social Media (2020.lrec-1)
Copied to clipboard
| Challenge: | Swearing is a common form of verbal communication and occurs in social media and online forums . a study by a team of researchers has investigated the phenomenon of swearing in Twitter . |
| Approach: | They analyze tweets to determine abusive swearing using models that automatically predict it . they also investigate lexical, syntactic, and affective features that are more informative . |
| Outcome: | The proposed model can predict abusive swearing in a tweet context and provide an intrinsic evaluation of the model. |
Multilingual Irony Detection with Dependency Syntax and Neural Models (2020.coling-main)
Copied to clipboard
Alessandra Teresa Cignarella, Valerio Basile, Manuela Sanguinetti, Cristina Bosco, Paolo Rosso, Farah Benamara
| Challenge: | Several semantic and syntactic devices can be used to express irony, causing the incongruity, determine the clash and play the role of irony triggers within a text. |
| Approach: | They propose to exploit linguistic resources where syntax is annotated according to the Universal Dependencies scheme. |
| Outcome: | The proposed method exploits linguistic resources where syntax is annotated according to the Universal Dependencies scheme. |
Quantifying the Influence of Irrelevant Contexts on Political Opinions Produced by LLMs (2025.acl-srw)
Copied to clipboard
| Challenge: | Recent studies have examined the generation of large language models (LLMs) on subjective topics such as political opinions and attitudinal questionnaires. |
| Approach: | They use a Political Compass Test questionnaire to quantify how irrelevant information can systematically bias model opinions in specific directions. |
| Outcome: | The results show that even seemingly unrelated contexts alter model responses in predictable ways. |
I Feel Offended, Don’t Be Abusive! Implicit/Explicit Messages in Offensive and Abusive Language (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent literature suggests different approaches to identify abusive language phenomena . however, there is a lack of data sets that take into account the degree of explicitness . |
| Approach: | They propose to use annotation guidelines to distinguish between explicit and implicit abuse in English and apply them to OLID/OffensEval. |
| Outcome: | The proposed tool distinguishes between explicit and implicit abuse in English and takes into account the degree of explicitness. |
APPReddit: a Corpus of Reddit Posts Annotated for Appraisal (2022.lrec-1)
Copied to clipboard
Marco Antonio Stranisci, Simona Frenda, Eleonora Ceccaldi, Valerio Basile, Rossana Damiano, Viviana Patti
| Challenge: | Existing resources for emotion recognition are lacking for appraisal models. |
| Approach: | They propose to use APPReddit to annotate non-experimental data according to Appraisal theories . they compare it with enISEAR, a corpus of events created in an experimental setting and annotated according to this theory. |
| Outcome: | The proposed model predicts four appraisal dimensions without significant loss . the proposed model is compared with enISEAR, a corpus of events created in an experimental setting and annotated for appraisal. |