Papers by Jorge Palomar-Giner

2 papers
Building a Data Infrastructure for a Mid-Resource Language: The Case of Catalan (2024.lrec-main)

Copied to clipboard

Challenge: Aina Project aims to provide Catalan with the resources needed to keep its relevance in AI/NLP applications.
Approach: They propose a set of strategies to consider when improving technology support for a mid- or low-resource language . they propose annotated datasets and a framework to make models ready to use .
Outcome: The Aina Project aims to provide Catalan with the necessary resources to keep its relevance in AI/NLP-related industry and research.
A CURATEd CATalog: Rethinking the Extraction of Pretraining Corpora for Mid-Resourced Languages (2024.lrec-main)

Copied to clipboard

Challenge: CATalog 1.0 is the largest text corpus in Catalan to date . CURATE is a pipeline that can be parallelizable to run in high performance clusters .
Approach: They propose a data pipeline that uses binary filters to filter documents based on text quality . they optimised the pipeline to run in high performance clusters .
Outcome: The proposed pipeline is optimized for high performance cluster environments and runs in high performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations