Papers by Elena Volodina
Superlim: A Swedish Language Understanding Evaluation Benchmark (2023.emnlp-main)
Copied to clipboard
Aleksandrs Berdicevskis, Gerlof Bouma, Robin Kurtz, Felix Morger, Joey Öhman, Yvonne Adesam, Lars Borin, Dana Dannélls, Markus Forsberg, Tim Isbister, Anna Lindahl, Martin Malmsten, Faton Rekathati, Magnus Sahlgren, Elena Volodina, Love Börjeson, Simon Hengchen, Nina Tahmasebi
| Challenge: | In this paper, we present a multi-task benchmark for Swedish language models . we address methodological challenges, such as mitigating the Anglocentric bias when creating datasets for a less-resourced language . |
| Approach: | They propose a multi-task NLP benchmark for Swedish language models . they propose to use superlim to evaluate Swedish language model performance . |
| Outcome: | The proposed benchmark does not approach ceiling performance on any of the tasks, suggesting it is difficult to implement. |
Towards Privacy by Design in Learner Corpora Research: A Case of On-the-fly Pseudonymization of Swedish Learner Essays (2020.coling-main)
Copied to clipboard
| Challenge: | An ongoing project aims at automating pseudonymization of learner essays . 89% of the personal information can be successfully identified in learner data . |
| Approach: | They propose to use rule-based methods to detect 15 categories out of 19 suggested by the authors. |
| Outcome: | The proposed methods detect 15 categories out of 19 suggested by the authors . 89% of the personal information can be successfully identified in learner data and annotated correctly with an inter-annotator agreement of 86% measured as Fleiss kappa and Krippendorff’s alpha. |
Pseudonymization Categories across Domain Boundaries (2024.lrec-main)
Copied to clipboard
Maria Irena Szawerna, Simon Dobnik, Therese Lindström Tiedemann, Ricardo Muñoz Sánchez, Xuan-Son Vu, Elena Volodina
| Challenge: | Linguistic data can contain personal information, which is limited in accessibility . a universal system of tags for categorizing PIIs could be developed to replace them . |
| Approach: | They analyze tagsets used for anonymization and pseudonymization to find out what kinds of PII appear in different domains. |
| Outcome: | The proposed system would allow for dynamic pseudonymization while keeping the data readable and useful for future research. |
Towards an Ideal Tool for Learner Error Annotation (2024.lrec-main)
Copied to clipboard
| Challenge: | 'correction annotation' is a technique that has been used for many years to correct errors in learner corpora. |
| Approach: | They propose to use SVALA to annotate and analyse corrections in learner corpora using a parallel aligned approach to visualisation and annotation. |
| Outcome: | The proposed tool supports multiple annotation systems, localisation into other languages, and the development of more complex annotation systems. |
UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment (2025.emnlp-main)
Copied to clipboard
Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Joshua Reynolds, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi
| Challenge: | Language proficiency research plays a central role in education and often intersects with advances in linguistics and AI. |
| Approach: | They propose a multilingual multidimensional dataset of texts annotated according to the CEFR scale in 13 languages. |
| Outcome: | The proposed dataset supports linguistic features and pretrained models in multilingual CEFR level assessment. |