Papers by Joan Llop

3 papers
FLOR: On the Effectiveness of Language Adaptation (2024.lrec-main)

Copied to clipboard

Challenge: Large language models have amply proven their capabilities, but low- and mid-resource languages do not have access to the necessary means to train such models from scratch.
Approach: They use a 26B tokens corpus to further pre-train BLOOM, giving rise to FLOR models.
Outcome: The proposed model achieves consistent gains across Catalan and Spanish tasks.
A weakly supervised textual entailment approach to zero-shot text classification (2023.eacl-main)

Copied to clipboard

Challenge: Existing methods to train on weakly supervised datasets are expensive due to the computational cost of pre-training.
Approach: They propose a method that trains on a weakly supervised dataset that is used as a proxy for a textual entailment problem and a target zero-shot text classification task.
Outcome: The proposed model achieves state-of-the-art performance in the scientific domain and competitive results in other areas.
A CURATEd CATalog: Rethinking the Extraction of Pretraining Corpora for Mid-Resourced Languages (2024.lrec-main)

Copied to clipboard

Challenge: CATalog 1.0 is the largest text corpus in Catalan to date . CURATE is a pipeline that can be parallelizable to run in high performance clusters .
Approach: They propose a data pipeline that uses binary filters to filter documents based on text quality . they optimised the pipeline to run in high performance clusters .
Outcome: The proposed pipeline is optimized for high performance cluster environments and runs in high performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations