Papers by Cristina Bosco

12 papers
Italian NLP for Everyone: Resources and Models from EVALITA to the European Language Grid (2022.lrec-1)

Copied to clipboard

Challenge: European Language Grid enables researchers and practitioners to easily distribute and use NLP resources and models.
Approach: They propose to integrate Italian NLP resources into the European Language Grid . they show how easy it is to use the integrated systems and demonstrate how seamless it is .
Outcome: The European Language Grid enables researchers and practitioners to easily distribute and use NLP resources and models.
PoSTWITA-UD: an Italian Twitter Treebank in Universal Dependencies (L18-1)

Copied to clipboard

Challenge: Various approaches and ad hoc resources are needed to provide proper coverage of specific linguistic phenomena.
Approach: They propose to annotate tweets using a well-known dependency-based annotation format . they propose to use the tweets for training NLP systems to improve their performance .
Outcome: The proposed resource can be used for training of NLP systems on social media texts.
Marking Irony Activators in a Universal Dependencies Treebank: The Case of an Italian Twitter Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing annotations for irony are difficult, and the recognition of it is difficult due to its polarity.
Approach: They propose a fine-grained annotation scheme centered on irony that highlights the tokens responsible for its activation and their morpho-syntactic features.
Outcome: The proposed scheme highlights the tokens responsible for irony activation and their morpho-syntactic features.
QUEEREOTYPES: A Multi-Source Italian Corpus of Stereotypes towards LGBTQIA+ Community Members (2024.lrec-main)

Copied to clipboard

Challenge: a dataset of social media texts addressing LGBTQIA+ individuals is presented in this paper . the dataset is based on two sources in italian: Facebook and Twitter .
Approach: They describe a dataset composed of two sub-corpora from two different sources in Italian . the dataset includes social media texts regarding LGBTQIA+ individuals, behaviors, ideology and events .
Outcome: The QUEEREOTYPES dataset includes social media texts regarding LGBTQIA+ individuals, behaviors, ideology and events.
An Italian Twitter Corpus of Hate Speech against Immigrants (L18-1)

Copied to clipboard

Challenge: a recent study has annotated 6,000 tweets for hate speech against immigrants . the annotation scheme was designed to account for the multiplicity of factors that can contribute to the definition of a hate speech notion .
Approach: They describe a Twitter corpus annotated for hate speech against immigrants . they propose a scheme that includes aggressiveness, offensiveness, irony, stereotype and intensity .
Outcome: The proposed annotation scheme includes aggressiveness, offensiveness, irony, stereotype, intensity and (on an experimental basis) intensity.
EPIC: Multi-Perspective Annotation of a Corpus of Irony (2023.acl-long)

Copied to clipboard

Challenge: EPIC is the first annotated corpus for irony analysis based on data perspectivism . a recent trend in natural language processing (NLP) postulates that the disagreement among annotators in a language resource is a valuable source of knowledge, rather than noise that ought to be minimized or discarded.
Approach: They propose to annotate an English perspectivist irony corpus based on data perspectivism . they validate the model by creating perspective-aware models that encode the perspectives of annotators grouped according to their demographic characteristics.
Outcome: The proposed model can capture different perspectives on irony among different groups of annotators, and is more confident than non-perspectivist models.
Treebanking User-Generated Content: A Proposal for a Unified Representation in Universal Dependencies (2020.lrec-1)

Copied to clipboard

Challenge: Despite the increasing number of contributions on Part-of-Speech tagging and parsing, automatic processing of user-generated content (UGC) still represents a challenging task.
Approach: They propose a set of guidelines for the annotation of user-generated texts within the Universal Dependencies framework.
Outcome: The proposed annotation guidelines promote cross-linguistic consistency, which has always been in the spirit of UD.
Confidence-based Ensembling of Perspective-aware Models (2023.emnlp-main)

Copied to clipboard

Challenge: Human label variability has been a topic of research in the field of NLP recently . Exploiting disagreements in annotations has been shown to offer advantages for accurate modelling and fairer evaluation.
Approach: They propose a highly perspectivist model that exploits disagreements in annotations to capture the subjectivity encoded in the annotation process.
Outcome: The proposed model is validated on irony and hate speech detection scenarios in in-domain and cross-domain settings.
UINAUIL: A Unified Benchmark for Italian Natural Language Understanding (2023.acl-demo)

Copied to clipboard

Challenge: a benchmark of six tasks for Italian Natural Language Understanding is presented . large language models (LLMs) have revolutionized the field of natural language processing . a few benchmarks exist for non-English languages, but only a handful are available for nonEnglish languages .
Approach: They introduce a benchmark for Italian Natural Language Understanding that harmonizes the data format and exposes functionalities to facilitate data manipulation and evaluation of custom models.
Outcome: The proposed benchmarks are based on the European Language Grid and available models in Italian and multilingual languages.
A Multilingual Dataset of Racial Stereotypes in Social Media Conversational Threads (2023.findings-eacl)

Copied to clipboard

Challenge: a new corpus-based study addresses racial stereotypes in social media conversations . a multilingual corpus of rhs is used to investigate how they are spread .
Approach: They propose a corpus-based method for multilingual racial stereotype identification in social media conversational threads.
Outcome: The proposed method sheds light on how racial hoaxes are spread and allows identification of negative stereotypes that reinforce them.
Application and Analysis of a Multi-layered Scheme for Irony on the Italian Twitter Corpus TWITTIRÒ (L18-1)

Copied to clipboard

Challenge: Using a multi-layered scheme for the fine-grained annotation of irony on Italian Twitter is a challenging task to be performed by both human annotators and automatic NLP systems.
Approach: They propose to apply a multi-layered scheme for the fine-grained annotation of irony to an Italian Twitter corpus.
Outcome: The proposed scheme can be validated on Italian irony-laden social media contents and is available in the cross- and multi-lingual perspective.
Multilingual Irony Detection with Dependency Syntax and Neural Models (2020.coling-main)

Copied to clipboard

Challenge: Several semantic and syntactic devices can be used to express irony, causing the incongruity, determine the clash and play the role of irony triggers within a text.
Approach: They propose to exploit linguistic resources where syntax is annotated according to the Universal Dependencies scheme.
Outcome: The proposed method exploits linguistic resources where syntax is annotated according to the Universal Dependencies scheme.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations