Papers by Viviana Cotik
MessIRve: A Large-Scale Spanish Information Retrieval Dataset (2025.emnlp-main)
Copied to clipboard
Francisco Valentini, Viviana Cotik, Damián Furman, Ivan Bercovich, Edgar Altszyler, Juan Manuel Pérez
| Challenge: | Information retrieval (IR) is the task of finding relevant documents in response to a user query. |
| Approach: | They propose a large-scale Spanish IR dataset with almost 700,000 queries from Google’s autocomplete API and relevant documents sourced from Wikipedia. |
| Outcome: | The proposed dataset covers a wide variety of topics, unlike smaller datasets. |
Indigenous Languages Spoken in Argentina: A Survey of NLP and Speech Resources (2025.coling-main)
Copied to clipboard
| Challenge: | Currently, no unified information on speakers and computational tools are available for these languages. |
| Approach: | They present a systematization of the indigenous languages spoken in Argentina, along with national demographic data on the country’s Indigenous population. |
| Outcome: | The proposed systematization of the indigenous languages spoken in Argentina, along with national demographic data on the country’s Indigenous population, is based on the Argentine population. |
Exploring Large Language Models for Hate Speech Detection in Rioplatense Spanish (2025.findings-naacl)
Copied to clipboard
| Challenge: | Hate speech detection deals with many language variants, slang, nuances, and cultural nuances. |
| Approach: | They propose to use large language models to detect hate speech in Rioplatense Spanish . they compare their results to those of a state-of-the-art BERT classifier . |
| Outcome: | The proposed models show lower precision than the state-of-the-art classifier, but are sensitive to highly nuanced cases. |