Papers by Cristina Marco

5 papers
Building Sentiment Lexicons for Mainland Scandinavian Languages Using Machine Translation and Sentence Embeddings (2022.lrec-1)

Copied to clipboard

Challenge: a simple but effective method to build sentiment lexicons for the three Mainland Scandinavian languages is proposed . a number of experiments with Scandinavian language datasets yield state-of-the-art results using a rule-based sentiment analysis algorithm.
Approach: They propose a simple but effective method to build sentiment lexicons for the three Mainland Scandinavian languages.
Outcome: The proposed method is based on the English Sentiwordnet and a thesaurus in one of the target languages.
An Italian Twitter Corpus of Hate Speech against Immigrants (L18-1)

Copied to clipboard

Challenge: a recent study has annotated 6,000 tweets for hate speech against immigrants . the annotation scheme was designed to account for the multiplicity of factors that can contribute to the definition of a hate speech notion .
Approach: They describe a Twitter corpus annotated for hate speech against immigrants . they propose a scheme that includes aggressiveness, offensiveness, irony, stereotype and intensity .
Outcome: The proposed annotation scheme includes aggressiveness, offensiveness, irony, stereotype, intensity and (on an experimental basis) intensity.
EPIC: Multi-Perspective Annotation of a Corpus of Irony (2023.acl-long)

Copied to clipboard

Challenge: EPIC is the first annotated corpus for irony analysis based on data perspectivism . a recent trend in natural language processing (NLP) postulates that the disagreement among annotators in a language resource is a valuable source of knowledge, rather than noise that ought to be minimized or discarded.
Approach: They propose to annotate an English perspectivist irony corpus based on data perspectivism . they validate the model by creating perspective-aware models that encode the perspectives of annotators grouped according to their demographic characteristics.
Outcome: The proposed model can capture different perspectives on irony among different groups of annotators, and is more confident than non-perspectivist models.
DGS-Fabeln-1: A Multi-Angle Parallel Corpus of Fairy Tales between German Sign Language and German Text (2024.lrec-main)

Copied to clipboard

Challenge: a parallel corpus of German text and videos containing fairy tales interpreted into the German Sign Language (DGS) is the first corpus filmed from 7 angles and one of the few sign language corpora globally which have been filmed simultaneously.
Approach: They present a parallel corpus of German fairy tales interpreted by a native DGS signer.
Outcome: The proposed corpus is the first semi-naturally expressed DGS that has been filmed from 7 angles and where the listener has been simultaneously filmed.
Semantic Diversity for Natural Language Understanding Evaluation in Dialog Systems (2020.coling-industry)

Copied to clipboard

Challenge: a dialog system is used to evaluate NLU models using aggregated metrics on a large number of utterances.
Approach: They propose a method to generate a test set with high semantic diversity for NLU evaluation in dialog systems.
Outcome: The proposed test sets are based on high diversity of utterances from different regions of the utteration embedding space.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations