Papers by Maite Melero

5 papers
Unmasking Biases: Exploring Gender Bias in English-Catalan Machine Translation through Tokenization Analysis and Novel Dataset (2024.lrec-main)

Copied to clipboard

Challenge: a new dataset focuses on gender-neutral terms that necessitate gendered translations in Catalan.
Approach: They propose to use a new dataset to evaluate gender bias in machine translation . they train four MT systems using different tokenization techniques .
Outcome: The proposed dataset focuses on gender-neutral terms necessitating gendered translations in Catalan.
On the Multilingual Capabilities of Very Large-Scale English Language Models (2022.lrec-1)

Copied to clipboard

Challenge: Generative Pre-trained Transformers (GPTs) have been scaled to unprecedented sizes in the history of machine learning.
Approach: They investigate the potential and limits of Generative Pre-trained Transformers in three tasks . they find it can be almost as useful for many languages as it is for English .
Outcome: The proposed model can perform tasks in five different languages, and its potential is explored . it can learn from a few examples "via text interaction" and is scalable to many languages .
Unsupervised Machine Translation in Real-World Scenarios (2022.lrec-1)

Copied to clipboard

Challenge: a recent study has shown that unsupervised methods rely on monolingual corpora to build MT systems.
Approach: They present the results of the MT4All CEF project using monolingual corpora . they propose to generate bilingual dictionaries and translation models from monolingual data .
Outcome: The proposed method generates bilingual dictionaries and translation models from monolingual corpora . results show that it is comparable to general domain supervised translation .
Spanish Datasets for Sensitive Entity Detection in the Legal Domain (2022.lrec-1)

Copied to clipboard

Challenge: The de-identification of sensible data is essential for data sharing and reuse, both for research and commercial purposes.
Approach: They propose to use four datasets annotated for named entity detection in Spanish to fine-tune models for the task of named entity-detection.
Outcome: The proposed model is based on four datasets annotated for named entity detection in Spanish with an estimated error rate of 14%.
Are Multilingual Models the Best Choice for Moderately Under-resourced Languages? A Comprehensive Assessment for Catalan (2021.findings-acl)

Copied to clipboard

Challenge: Multilingual language models have been a crucial breakthrough for under-resourced languages . however, the superiority of language-specific models has already been proven for underresourced ones .
Approach: They propose to build a monolingual monolingual model that is comparable to state-of-the-art large multilingual models.
Outcome: The proposed model consistently outperforms state-of-the-art models across tasks and settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations