Papers by Jorge Palomar-Giner
Building a Data Infrastructure for a Mid-Resource Language: The Case of Catalan (2024.lrec-main)
Copied to clipboard
Aitor Gonzalez-Agirre, Montserrat Marimon, Carlos Rodriguez-Penagos, Javier Aula-Blasco, Irene Baucells, Carme Armentano-Oller, Jorge Palomar-Giner, Baybars Kulebi, Marta Villegas
| Challenge: | Aina Project aims to provide Catalan with the resources needed to keep its relevance in AI/NLP applications. |
| Approach: | They propose a set of strategies to consider when improving technology support for a mid- or low-resource language . they propose annotated datasets and a framework to make models ready to use . |
| Outcome: | The Aina Project aims to provide Catalan with the necessary resources to keep its relevance in AI/NLP-related industry and research. |
A CURATEd CATalog: Rethinking the Extraction of Pretraining Corpora for Mid-Resourced Languages (2024.lrec-main)
Copied to clipboard
Jorge Palomar-Giner, Jose Javier Saiz, Ferran Espuña, Mario Mina, Severino Da Dalt, Joan Llop, Malte Ostendorff, Pedro Ortiz Suarez, Georg Rehm, Aitor Gonzalez-Agirre, Marta Villegas
| Challenge: | CATalog 1.0 is the largest text corpus in Catalan to date . CURATE is a pipeline that can be parallelizable to run in high performance clusters . |
| Approach: | They propose a data pipeline that uses binary filters to filter documents based on text quality . they optimised the pipeline to run in high performance clusters . |
| Outcome: | The proposed pipeline is optimized for high performance cluster environments and runs in high performance. |