Papers by Elena Álvarez-Mellado
ConLoan: A Contrastive Multilingual Dataset for Evaluating Loanwords (2025.acl-long)
Copied to clipboard
Sina Ahmadi, Micha David Hess, Elena Álvarez-Mellado, Alessia Battisti, Cui Ding, Anne Göhring, Yingqiang Gao, Zifan Jiang, Andrianos Michail, Peshmerge Morad, Joel Niklaus, Maria Christina Panagiotopoulou, Stefano Perrella, Juri Opitz, Anastassia Shaitarova, Rico Sennrich
| Challenge: | Lexical borrowing is a ubiquitous linguistic phenomenon influenced by geopolitical, societal, and technological factors. |
| Approach: | They propose a novel contrastive dataset comprising sentences with and without loanwords across 10 languages to examine how machine translation and language models process loanword . |
| Outcome: | The proposed dataset shows that state-of-the-art models prefer loanwords over native terms and exhibit varying performance across languages. |
A Corpus of Spanish Political Speeches from 1937 to 2019 (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of political speeches in Spanish is documented from 1937 to 2019 . the corpus contains the speeches delivered by the head of state of Spain on Christmas Eve . |
| Approach: | They propose to collect political speeches from the Christmas Eve national speeches from 1937 to 2019 . they propose a Python interface that allows querying and analyzing the corpus . |
| Outcome: | The proposed corpus contains speeches delivered by the king of Spain from 1937 to 2019 . the documents reflect some of the most significant events and political changes in recent history . a set of HTML visualizations is provided to navigate the corpus and explore differences between TF-IDF frequencies. |
Evaluating Sequence Labeling on the basis of Information Theory (2025.acl-long)
Copied to clipboard
| Challenge: | Existing metric families focus on certain aspects of sequence labeling tasks. |
| Approach: | They propose a metric that measures how much information each token contributes depending on different aspects of the sequence. |
| Outcome: | The proposed metric can satisfy all properties simultaneously. |
Detecting Unassimilated Borrowings in Spanish: An Annotated Corpus and Approaches to Modeling (2022.acl-long)
Copied to clipboard
| Challenge: | a corpus of Spanish newswire rich in unassimilated lexical borrowings is used to identify the language of a word. |
| Approach: | They propose to annotate a corpus of Spanish newswire rich in unassimilated lexical borrowings and evaluate how models perform on this task. |
| Outcome: | The proposed model outperforms models fed with subword embeddings and Transformer-based embeddables on the Spanish newswire corpus. |