Papers by Valerio Basile

16 papers
WebNLG-IT: Construction of an aligned RDF-Italian corpus through Machine Translation techniques (2025.findings-acl)

Copied to clipboard

Challenge: Using NMT and hand-written rules, we created the first aligned Italian RDF-to-text corpus .
Approach: They propose to use NMT to create an Italian version of the WebNLG corpus and to refine and improve the quality of the produced resource.
Outcome: The proposed system is the best on the original English version and the best in the second step, it improves and refines the quality of the produced resource.
The DipInfoUniTo Realizer at SRST’19: Learning to Rank and Deep Morphology Prediction for Multilingual Surface Realization (D19-63)

Copied to clipboard

Challenge: SR is one of the main tasks involved in Natural Language Generation.
Approach: They propose a system which divides the SR task into two independent subtasks, namely word order prediction and morphology inflection prediction.
Outcome: The proposed system is a direct successor to the architecture presented at SR'19.
Italian NLP for Everyone: Resources and Models from EVALITA to the European Language Grid (2022.lrec-1)

Copied to clipboard

Challenge: European Language Grid enables researchers and practitioners to easily distribute and use NLP resources and models.
Approach: They propose to integrate Italian NLP resources into the European Language Grid . they show how easy it is to use the integrated systems and demonstrate how seamless it is .
Outcome: The European Language Grid enables researchers and practitioners to easily distribute and use NLP resources and models.
PERSEVAL: A Framework for Perspectivist Classification Evaluation (2025.emnlp-main)

Copied to clipboard

Challenge: Perspectivist evaluation practices in NLP remain fragmented and inconsistent .
Approach: They propose a framework that evaluates perspectivist models at the individual annotator level and treats annotators and users as distinct entities, consistent with real-world scenarios.
Outcome: The proposed framework evaluates annotators and users as distinct entities consistent with real-world scenarios.
Capturing Perspectives of Crowdsourced Annotators in Subjective Learning Tasks (2024.naacl-long)

Copied to clipboard

Challenge: Existing approaches to label aggregation fail to capture subjective annotations and can lead to biases.
Approach: They propose annotator-aware representations for text for subjective classification tasks that involve learning representations of annotators.
Outcome: The proposed model improves on metrics that assess the performance on capturing individual annotators’ perspectives.
Label Augmentation for Zero-Shot Hierarchical Text Classification (2024.acl-long)

Copied to clipboard

Challenge: Hierarchical Text Classification is a difficult problem due to the lack of labeled data and the cost of manually annotating data samples.
Approach: They propose a method that uses a Large Language Model to augment the deepest layer of the labels hierarchy to enhance its specificity.
Outcome: The proposed method achieves state-of-the-art on four public datasets and a strong correlation between the metric values and the classification performance.
QUEEREOTYPES: A Multi-Source Italian Corpus of Stereotypes towards LGBTQIA+ Community Members (2024.lrec-main)

Copied to clipboard

Challenge: a dataset of social media texts addressing LGBTQIA+ individuals is presented in this paper . the dataset is based on two sources in italian: Facebook and Twitter .
Approach: They describe a dataset composed of two sub-corpora from two different sources in Italian . the dataset includes social media texts regarding LGBTQIA+ individuals, behaviors, ideology and events .
Outcome: The QUEEREOTYPES dataset includes social media texts regarding LGBTQIA+ individuals, behaviors, ideology and events.
I’m sure you’re a real scholar yourself: Exploring Ironic Content Generation by Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Moreover, irony is highly subjective and can depend on various factors, such as social, cultural, or generational aspects.
Approach: They propose to fine-tune two large language models to generate ironic and non-ironic content and analyze their outputs from a linguistic perspective.
Outcome: The proposed models generate ironic and non-ironic responses to a given social media post and analyze their outputs from a linguistic perspective.
EPIC: Multi-Perspective Annotation of a Corpus of Irony (2023.acl-long)

Copied to clipboard

Challenge: EPIC is the first annotated corpus for irony analysis based on data perspectivism . a recent trend in natural language processing (NLP) postulates that the disagreement among annotators in a language resource is a valuable source of knowledge, rather than noise that ought to be minimized or discarded.
Approach: They propose to annotate an English perspectivist irony corpus based on data perspectivism . they validate the model by creating perspective-aware models that encode the perspectives of annotators grouped according to their demographic characteristics.
Outcome: The proposed model can capture different perspectives on irony among different groups of annotators, and is more confident than non-perspectivist models.
Confidence-based Ensembling of Perspective-aware Models (2023.emnlp-main)

Copied to clipboard

Challenge: Human label variability has been a topic of research in the field of NLP recently . Exploiting disagreements in annotations has been shown to offer advantages for accurate modelling and fairer evaluation.
Approach: They propose a highly perspectivist model that exploits disagreements in annotations to capture the subjectivity encoded in the annotation process.
Outcome: The proposed model is validated on irony and hate speech detection scenarios in in-domain and cross-domain settings.
UINAUIL: A Unified Benchmark for Italian Natural Language Understanding (2023.acl-demo)

Copied to clipboard

Challenge: a benchmark of six tasks for Italian Natural Language Understanding is presented . large language models (LLMs) have revolutionized the field of natural language processing . a few benchmarks exist for non-English languages, but only a handful are available for nonEnglish languages .
Approach: They introduce a benchmark for Italian Natural Language Understanding that harmonizes the data format and exposes functionalities to facilitate data manipulation and evaluation of custom models.
Outcome: The proposed benchmarks are based on the European Language Grid and available models in Italian and multilingual languages.
Do You Really Want to Hurt Me? Predicting Abusive Swearing in Social Media (2020.lrec-1)

Copied to clipboard

Challenge: Swearing is a common form of verbal communication and occurs in social media and online forums . a study by a team of researchers has investigated the phenomenon of swearing in Twitter .
Approach: They analyze tweets to determine abusive swearing using models that automatically predict it . they also investigate lexical, syntactic, and affective features that are more informative .
Outcome: The proposed model can predict abusive swearing in a tweet context and provide an intrinsic evaluation of the model.
Multilingual Irony Detection with Dependency Syntax and Neural Models (2020.coling-main)

Copied to clipboard

Challenge: Several semantic and syntactic devices can be used to express irony, causing the incongruity, determine the clash and play the role of irony triggers within a text.
Approach: They propose to exploit linguistic resources where syntax is annotated according to the Universal Dependencies scheme.
Outcome: The proposed method exploits linguistic resources where syntax is annotated according to the Universal Dependencies scheme.
Quantifying the Influence of Irrelevant Contexts on Political Opinions Produced by LLMs (2025.acl-srw)

Copied to clipboard

Challenge: Recent studies have examined the generation of large language models (LLMs) on subjective topics such as political opinions and attitudinal questionnaires.
Approach: They use a Political Compass Test questionnaire to quantify how irrelevant information can systematically bias model opinions in specific directions.
Outcome: The results show that even seemingly unrelated contexts alter model responses in predictable ways.
I Feel Offended, Don’t Be Abusive! Implicit/Explicit Messages in Offensive and Abusive Language (2020.lrec-1)

Copied to clipboard

Challenge: Recent literature suggests different approaches to identify abusive language phenomena . however, there is a lack of data sets that take into account the degree of explicitness .
Approach: They propose to use annotation guidelines to distinguish between explicit and implicit abuse in English and apply them to OLID/OffensEval.
Outcome: The proposed tool distinguishes between explicit and implicit abuse in English and takes into account the degree of explicitness.
APPReddit: a Corpus of Reddit Posts Annotated for Appraisal (2022.lrec-1)

Copied to clipboard

Challenge: Existing resources for emotion recognition are lacking for appraisal models.
Approach: They propose to use APPReddit to annotate non-experimental data according to Appraisal theories . they compare it with enISEAR, a corpus of events created in an experimental setting and annotated according to this theory.
Outcome: The proposed model predicts four appraisal dimensions without significant loss . the proposed model is compared with enISEAR, a corpus of events created in an experimental setting and annotated for appraisal.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations