Papers by Irene Baucells

5 papers
Building a Data Infrastructure for a Mid-Resource Language: The Case of Catalan (2024.lrec-main)

Copied to clipboard

Challenge: Aina Project aims to provide Catalan with the resources needed to keep its relevance in AI/NLP applications.
Approach: They propose a set of strategies to consider when improving technology support for a mid- or low-resource language . they propose annotated datasets and a framework to make models ready to use .
Outcome: The Aina Project aims to provide Catalan with the necessary resources to keep its relevance in AI/NLP-related industry and research.
Dynamic Stance: Modeling Discussions by Labeling the Interactions (2023.findings-emnlp)

Copied to clipboard

Challenge: Stance detection is a popular task that has been modeled as a static task, but its limitations are strong topic-dependent.
Approach: They propose to model stance as a dynamic task by focusing on interactions between a message and their replies.
Outcome: The proposed model shows portability across topics and languages.
FLOR: On the Effectiveness of Language Adaptation (2024.lrec-main)

Copied to clipboard

Challenge: Large language models have amply proven their capabilities, but low- and mid-resource languages do not have access to the necessary means to train such models from scratch.
Approach: They use a 26B tokens corpus to further pre-train BLOOM, giving rise to FLOR models.
Outcome: The proposed model achieves consistent gains across Catalan and Spanish tasks.
Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages? (2025.emnlp-main)

Copied to clipboard

Challenge: a recent study focused on complex, high-level tasks, but LMentry is limited to English . a multilingual evaluation of large language models is needed to address this gap, authors say .
Approach: They propose a compact benchmark that enables systematic evaluation of large language models . they propose to use tasks that are trivial for humans but remain surprisingly difficult for LLMs .
Outcome: The proposed benchmark is limited to English, leaving its insights linguistically narrow.
IberoBench: A Benchmark for LLM Evaluation in Iberian Languages (2025.coling-main)

Copied to clipboard

Challenge: Existing multi-task benchmarks for Large Language Models are limited to English . a new benchmark is needed to evaluate models on a range of tasks .
Approach: They propose a multilingual, multi-task benchmark for Iberian languages built on the LM Evaluation Harness framework.
Outcome: The proposed benchmark covers 62 tasks divided into 179 subtasks and is available in Iberian, Basque, Catalan, Galician, European Spanish and European Portuguese.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations