Papers by Olga Pelloni
A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets (2024.findings-naacl)
Copied to clipboard
| Challenge: | a new study aims to assess linguistic diversity of multilingual data sets against a reference language sample . linguistic diversity is typically measured as the number of languages included in the data set . but such measures do not consider structural properties of the included languages . |
| Approach: | They propose to measure linguistic diversity against a reference language sample to maximise linguistic diversity. |
| Outcome: | The proposed measure can be used to identify the types of languages that are not represented in a data set. |
Subword Evenness (SuE) as a Predictor of Cross-lingual Transfer to Low-resource Languages (2022.emnlp-main)
Copied to clipboard
| Challenge: | English is the most natural choice for cross-lingual transfer, but it is often not the best choice for low-resource languages. |
| Approach: | They propose to use pre-trained multilingual models to improve performance in low-resource languages via cross-lingual transfer. |
| Outcome: | The results show that languages written in non-Latin and non-alphabetic scripts are the best choices for improving performance on Masked Language Modelling tasks in a diverse set of 30 low-resource languages. |
TeDDi Sample: Text Data Diversity Sample for Language Comparison and Multilingual NLP (2022.lrec-1)
Copied to clipboard
| Challenge: | a deep understanding of language is being achieved through increased access to data from minority and low-resource languages. |
| Approach: | They present a diversity sample of text data for language comparison and multilingual natural language processing. |
| Outcome: | The TeDDi sample features 89 languages based on the typological diversity sample in the World Atlas of Language Structures . |