Papers by Luis Chiruzzo
Development of a Guarani - Spanish Parallel Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of Guarani sentences with sentence-level alignment is presented . the corpus contains 228,000 Guaran tokens along with 336,000 Spanish tokens . |
| Approach: | They propose to develop a Guarani - Spanish parallel corpus with sentence-level alignment . the corpus contains 228,000 Guaran tokens along with 336,000 Spanish tokens . |
| Outcome: | The proposed corpus contains 22,800 Guarani tokens along with 336,000 Spanish tokens extracted from web sources. |
Jojajovai: A Parallel Guarani-Spanish Corpus for MT Benchmarking (2022.lrec-1)
Copied to clipboard
Luis Chiruzzo, Santiago Góngora, Aldo Alvarez, Gustavo Giménez-Lugo, Marvin Agüero-Torales, Yliana Rodríguez
| Challenge: | a corpus of Guarani-Spanish text is presented that is aligned at sentence level . the long history of language contact between Guaran and Spanish in South America has resulted in many interesting language varieties . |
| Approach: | They propose to align Guarani-Spanish text at sentence level with 30,000 sentence pairs and a test set. |
| Outcome: | The proposed corpus contains about 30,000 sentence pairs and is structured as a collection of subsets from different sources, further split into training, development and test sets. |
Grammar-based Data Augmentation for Low-Resource Languages: The Case of Guarani-Spanish Neural Machine Translation (2024.naacl-long)
Copied to clipboard
Agustín Lucas, Alexis Baladón, Victoria Pardiñas, Marvin Agüero-Torales, Santiago Góngora, Luis Chiruzzo
| Challenge: | Low-resource languages suffer from a vicious circle: data is needed to build tools, but available text is scarce. |
| Approach: | They propose to use a grammar-based system to generate Spanish text and syntactically transfer it to Guarani to boost its performance. |
| Outcome: | The proposed system outperforms existing models by pretraining models with synthetic text. |
AmericasNLI: Evaluating Zero-shot Natural Language Understanding of Pretrained Multilingual Models in Truly Low-resource Languages (2022.acl-long)
Copied to clipboard
Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir Meza Ruiz, Gustavo Giménez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Thang Vu, Katharina Kann
| Challenge: | Pretrained multilingual models can perform cross-lingual transfer in a zero-shot setting, even for unseen languages. |
| Approach: | They propose to extend XNLI to 10 indigenous languages of the Americas and test multiple zero-shot and translation-based approaches. |
| Outcome: | The proposed model can perform cross-lingual transfer in a zero-shot setting even for languages unseen during pretraining. |
HAHA 2019 Dataset: A Corpus for Humor Analysis in Spanish (2020.lrec-1)
Copied to clipboard
| Challenge: | 30,000 Spanish tweets were crowd-annotated with humor value and funniness score . the corpus contains approximately 38.6% of humorous tweets with an average score of 2.04 in a scale from 1 to 5 for the humorous tweet. |
| Approach: | They develop a corpus of 30,000 Spanish tweets crowd-annotated with humor value and funniness score. |
| Outcome: | The results obtained from the 30,000 tweets in the Spanish language are encouraging. |
La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America (2025.acl-long)
Copied to clipboard
María Grandury, Javier Aula-Blasco, Júlia Falcão, Clémentine Fourrier, Miguel González Saiz, Gonzalo Martínez, Gonzalo Santamaria Gomez, Rodrigo Agerri, Nuria Aldama García, Luis Chiruzzo, Javier Conde, Helena Gomez Adorno, Marta Guerrero Nieto, Guido Ivetta, Natàlia López Fuertes, Flor Miriam Plaza-del-Arco, María-Teresa Martín-Valdivia, Helena Montoro Zamorano, Carmen Muñoz Sanz, Pedro Reviriego, Leire Rosado Plaza, Alejandro Vaca Serrano, Estrella Vallecillo-Rodríguez, Jorge Vallego, Irune Zubiaga
| Challenge: | La Leaderboard is the first open-source leaderboard to evaluate generative Large Language Models (LLMs) in languages and language varieties of Spain and Latin America. |
| Approach: | They propose to use La Leaderboard to evaluate generative Large Language Models in Spanish and Latin America. |
| Outcome: | La Leaderboard is the first open-source leaderboard to evaluate generative LLMs in languages and language varieties of Spain and Latin America. |
Spanish HPSG Treebank based on the AnCora Corpus (L18-1)
Copied to clipboard
| Challenge: | a corpus of HPSG annotated trees for Spanish contains morphosyntactic information, annotations for semantic roles, clitic pronouns and relative clauses. |
| Approach: | They propose to build a Spanish HPSG annotated corpus based on the Spanish corpus AnCora and an HTML format for visualizing the trees in a browser. |
| Outcome: | The proposed corpus contains syntactic and morphological information, semantic roles, clitic pronouns and relative clauses, and has CFG style annotations. |
A Multi-level Annotated Corpus of Scientific Papers for Scientific Document Summarization and Cross-document Relation Discovery (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent studies have proposed to take advantage of the scientific paper's citation network to approach literature summarization. |
| Approach: | They propose to annotate related work sections, cite papers and sentences using machine readable data and an additional layer of papers citing the references. |
| Outcome: | The proposed corpus expands the existing data-set of related work sections and cites the papers cited in the related work section. |
Meeting the Needs of Low-Resource Languages: The Value of Automatic Alignments via Pretrained Models (2023.eacl-main)
Copied to clipboard
Abteen Ebrahimi, Arya D. McCarthy, Arturo Oncevay, John E. Ortega, Luis Chiruzzo, Gustavo Giménez-Lugo, Rolando Coto-Solano, Katharina Kann
| Challenge: | Large multilingual models have inspired a new class of word alignment methods, which work well for pretraining languages. |
| Approach: | They propose to use transformer-based word alignment methods to extract alignments from massive pretrained models. |
| Outcome: | The proposed methods outperform traditional methods for languages unseen to pretraining models, and are competitive with each other. |
Null Subjects in Spanish as a Machine Translation Problem (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for detecting null subjects and impersonal constructions in Spanish are limited. |
| Approach: | They adapt a machine translation methodology to detect null subjects and impersonal constructions in Spanish using an AnCora corpus. |
| Outcome: | The proposed approach surpasses the state-of-the-art in the detection of null subjects and impersonal constructions in Spanish while using modest computational resources. |