Papers by Carme Armentano-Oller

3 papers
Building a Data Infrastructure for a Mid-Resource Language: The Case of Catalan (2024.lrec-main)

Copied to clipboard

Challenge: Aina Project aims to provide Catalan with the resources needed to keep its relevance in AI/NLP applications.
Approach: They propose a set of strategies to consider when improving technology support for a mid- or low-resource language . they propose annotated datasets and a framework to make models ready to use .
Outcome: The Aina Project aims to provide Catalan with the necessary resources to keep its relevance in AI/NLP-related industry and research.
Becoming a High-Resource Language in Speech: The Catalan Case in the Common Voice Corpus (2024.lrec-main)

Copied to clipboard

Challenge: a project to create a publicly available voice dataset for speech recognition systems in Catalan is a multifaceted challenge.
Approach: They propose to create a publicly available voice dataset for future speech technologies in Catalan using the Mozilla Common Voice crowd-sourcing platform.
Outcome: The proposed dataset shows that Catalan ranks as the most prominent language in the corpus.
Are Multilingual Models the Best Choice for Moderately Under-resourced Languages? A Comprehensive Assessment for Catalan (2021.findings-acl)

Copied to clipboard

Challenge: Multilingual language models have been a crucial breakthrough for under-resourced languages . however, the superiority of language-specific models has already been proven for underresourced ones .
Approach: They propose to build a monolingual monolingual model that is comparable to state-of-the-art large multilingual models.
Outcome: The proposed model consistently outperforms state-of-the-art models across tasks and settings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations