Papers by Baiba Saulite
Creation of a Balanced State-of-the-Art Multilayer Corpus for NLU (L18-1)
Copied to clipboard
Normunds Gruzitis, Lauma Pretkalnina, Baiba Saulite, Laura Rituma, Gunta Nespore-Berzkalne, Arturs Znotins, Peteris Paikens
| Challenge: | Using full stack of language resources, we are creating a balanced text corpus for Latvian. |
| Approach: | They propose to create a syntactically and semantically annotated multilayered corpus for Latvian . they use widely acknowledged and cross-lingual representations for the corpus . |
| Outcome: | The proposed corpus adopts widely recognized and cross-lingual representations for natural language understanding and generation in Latvian. |
Latvian National Corpora Collection – Korpuss.lv (2022.lrec-1)
Copied to clipboard
Baiba Saulite, Roberts Darģis, Normunds Gruzitis, Ilze Auzina, Kristīne Levāne-Petrova, Lauma Pretkalniņa, Laura Rituma, Peteris Paikens, Arturs Znotins, Laine Strankale, Kristīne Pokratniece, Ilmārs Poikāns, Guntis Barzdins, Inguna Skadiņa, Anda Baklāne, Valdis Saulespurēns, Jānis Ziediņš
| Challenge: | Latvian National Corpora Collection (LNCC) is a multi-institutional and multi-project effort supporting the Latvian language research and language modelling. |
| Approach: | They propose to use Latvian corpora for linguistic research and language modelling. |
| Outcome: | LNCC is a multi-institutional and multi-project effort supported by the Digital Humanities and Language Technology communities in Latvia. |
BalsuTalka.lv - Boosting the Common Voice Corpus for Low-Resource Languages (2024.lrec-main)
Copied to clipboard
Roberts Dargis, Arturs Znotins, Ilze Auzina, Baiba Saulite, Sanita Reinsone, Raivis Dejus, Antra Klavinska, Normunds Gruzitis
| Challenge: | Latvian is a low-resource language for many NLP tasks, but most speech corpora are closed data . a crowdsourcing campaign to create a relatively large, diverse and open speech corpus for Latvian has been launched . |
| Approach: | a crowdsourcing campaign is helping to create an open speech corpus for Latvian . the goal is to enlarge the datasets and make them more diverse . authors use the opensource Mozilla Common Voice platform to validate speech samples . |
| Outcome: | a crowdsourcing initiative has increased the size and speaker diversity of the Latvian Common Voice 17.0 dataset by more than tenfold in less than a year. |