Papers by Pablo Gamallo
Truth Knows No Language: Evaluating Truthfulness Beyond English (2025.acl-long)
Copied to clipboard
Blanca Calvo Figueras, Eneko Sagarzazu, Julen Etxaniz, Jeremy Barnes, Pablo Gamallo, Iria de-Dios-Flores, Rodrigo Agerri
| Challenge: | a new benchmark evaluates the truthfulness of large language models (LLMs) based on imitative falsehoods. |
| Approach: | They propose a professionally translated extension of the TruthfulQA benchmark . it evaluates truthfulness in Basque, Catalan, Galician, and Spanish . |
| Outcome: | The proposed extension of the TruthfulQA benchmark evaluates truthfulness in Basque, Catalan, Galician, and Spanish. |
PartisanLens: A Multilingual Dataset of Hyperpartisan and Conspiratorial Immigration Narratives in European Media (2026.eacl-long)
Copied to clipboard
Michele Joshua Maggini, Paloma Piot, Anxo Pérez, Erik Bran Marino, Lúa Santamaría Montesinos, Ana Lisboa Cotovio, Marta Vázquez Abuín, Javier Parapar, Pablo Gamallo
| Challenge: | Existing methods for detecting hyperpartisan narratives and PRCTs are limited . hyperpartisan content promotes extreme views through one-sided, emotional language . |
| Approach: | They propose a multilingual dataset of 1617 hyperpartisan news headlines in Spanish, Italian, and Portuguese annotated in multiple political discourse aspects. |
| Outcome: | The proposed dataset is the first multilingual dataset of 1617 hyperpartisan headlines in Spanish, Italian, and Portuguese. |
Par-ITA: Benchmarking Seq2Seq and LLMs on a Human-Supervised Parallel Corpus for Italian Hyperpartisan Neutralization (2026.acl-long)
Copied to clipboard
| Challenge: | a new study examines the role of hyperpartisan content in online polarization in the social web. |
| Approach: | They propose a human-supervised parallel corpus for italian hyperpartisan neutralization of 2,475 paragraph pairs. |
| Outcome: | The proposed dataset is the first human-supervised parallel corpus for italian hyperpartisan neutralization of 2,475 paragraph pairs. |
Continued Pretraining and Interpretability-Based Evaluation for Low-Resource Languages: A Galician Case Study (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in large language models have led to remarkable improvements in language understanding and text generation. |
| Approach: | They propose a framework to evaluate large language models for underrepresented languages . they examine CPT strategies for languages with limited representation in multilingual models . |
| Outcome: | The proposed evaluation framework is based on the case of Galician language . it assesses trade-offs between linguistic enrichment and task-solving capabilities . |
IberoBench: A Benchmark for LLM Evaluation in Iberian Languages (2025.coling-main)
Copied to clipboard
Irene Baucells, Javier Aula-Blasco, Iria de-Dios-Flores, Silvia Paniagua Suárez, Naiara Perez, Anna Salles, Susana Sotelo Docio, Júlia Falcão, Jose Javier Saiz, Robiert Sepulveda Torres, Jeremy Barnes, Pablo Gamallo, Aitor Gonzalez-Agirre, German Rigau, Marta Villegas
| Challenge: | Existing multi-task benchmarks for Large Language Models are limited to English . a new benchmark is needed to evaluate models on a range of tasks . |
| Approach: | They propose a multilingual, multi-task benchmark for Iberian languages built on the LM Evaluation Harness framework. |
| Outcome: | The proposed benchmark covers 62 tasks divided into 179 subtasks and is available in Iberian, Basque, Catalan, Galician, European Spanish and European Portuguese. |