Papers by Júlia Falcão

6 papers
Cognitive Biases, Task Complexity, and Result Intepretability in Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Recent work shows that cognitive biases occur frequently in language models . a cognitive bias is a systematic deviation in judgment that simplifies complex decisions .
Approach: They evaluate the performance of different groups of models for each type of cognitive bias . they find that task complexity plays a part in eliciting stronger effects for some biases .
Outcome: The proposed models perform better for each type of bias in different settings . the results show that task complexity plays a part in eliciting stronger effects .
COMET for Low-Resource Machine Translation Evaluation: A Case Study of English-Maltese and Spanish-Basque (2024.lrec-main)

Copied to clipboard

Challenge: Trainable metrics for machine translation evaluation have been scoring the highest correlations with human judgements in the meta-evaluations.
Approach: They run a crowd-based evaluation campaign to evaluate COMET-22 and fine-tune it to improve its performance.
Outcome: The proposed system outperforms BLEU and other lexical overlap metrics in the meta-evaluations.
VeritasQA: A Truthfulness Benchmark Aimed at Multilingual Transferability (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) struggle with falsehoods and model hallucination . many efforts struggle to surpass 50% accuracy, with only targeted techniques reaching around 65% .
Approach: They propose a truthfulness benchmark that focuses on imitative falsehoods . they use a set of 353 questions and answers inspired by common misconceptions based on the language .
Outcome: The benchmark is available in Spanish, Catalan, Galician and English . it measures the truthfulness of multilingual LLMs using 353 questions and answers .
La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America (2025.acl-long)

Copied to clipboard

Challenge: La Leaderboard is the first open-source leaderboard to evaluate generative Large Language Models (LLMs) in languages and language varieties of Spain and Latin America.
Approach: They propose to use La Leaderboard to evaluate generative Large Language Models in Spanish and Latin America.
Outcome: La Leaderboard is the first open-source leaderboard to evaluate generative LLMs in languages and language varieties of Spain and Latin America.
Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages? (2025.emnlp-main)

Copied to clipboard

Challenge: a recent study focused on complex, high-level tasks, but LMentry is limited to English . a multilingual evaluation of large language models is needed to address this gap, authors say .
Approach: They propose a compact benchmark that enables systematic evaluation of large language models . they propose to use tasks that are trivial for humans but remain surprisingly difficult for LLMs .
Outcome: The proposed benchmark is limited to English, leaving its insights linguistically narrow.
IberoBench: A Benchmark for LLM Evaluation in Iberian Languages (2025.coling-main)

Copied to clipboard

Challenge: Existing multi-task benchmarks for Large Language Models are limited to English . a new benchmark is needed to evaluate models on a range of tasks .
Approach: They propose a multilingual, multi-task benchmark for Iberian languages built on the LM Evaluation Harness framework.
Outcome: The proposed benchmark covers 62 tasks divided into 179 subtasks and is available in Iberian, Basque, Catalan, Galician, European Spanish and European Portuguese.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations