Papers by Kaja Dobrovoljc
Gigafida 2.0: The Reference Corpus of Written Standard Slovene (2020.lrec-1)
Copied to clipboard
Simon Krek, Špela Arhar Holdt, Tomaž Erjavec, Jaka Čibej, Andraz Repar, Polona Gantar, Nikola Ljubešić, Iztok Kosem, Kaja Dobrovoljc
| Challenge: | Gigafida reference corpus of Slovene is updated with new material and tools . focus of upgrade was on transformation from general reference corp to standard reference corp . |
| Approach: | We present a new version of the Gigafida reference corpus of Slovene . the upgrade includes new material and better tools for annotating it . |
| Outcome: | The new version of the Gigafida reference corpus of Slovene is described . the whole Gigido corpus was deduplicated for the first time . |
Spoken Language Treebanks in Universal Dependencies: an Overview (2022.lrec-1)
Copied to clipboard
| Challenge: | spoken language treebanks have divergent annotation schemes limiting cross-resource explorations . many spoken language trees have no written form, but many of the world languages have no spoken form at all. |
| Approach: | They propose to use the Universal Dependencies annotation scheme to annotate spoken language treebanks using a morphosyntactic annotation scheme. |
| Outcome: | The proposed treebanks differ significantly with respect to the inventory and format of transcribed phenomena and the principles adopted in their morphosyntactic annotation. |
Gos 2: A New Reference Corpus of Spoken Slovenian (2024.lrec-main)
Copied to clipboard
| Challenge: | a new corpus of spoken Slovenian has been added to the Gos reference corpus . the corpus is now more than double the original size of 300 hours, 2.4 million words . |
| Approach: | They propose to add speech recordings and transcriptions from two related initiatives, the Gos VideoLectures corpus of public academic speech, and the Artur speech recognition database. |
| Outcome: | The new corpus is double the original size and contains 2.4 million words . it includes speech recordings and transcriptions from two related initiatives . |
SUK 1.0: A New Training Corpus for Linguistic Annotation of Modern Standard Slovene (2024.lrec-main)
Copied to clipboard
Špela Arhar Holdt, Jaka Čibej, Kaja Dobrovoljc, Tomaž Erjavec, Polona Gantar, Simon Krek, Tina Munda, Nejc Robida, Luka Terčon, Slavko Zitnik
| Challenge: | a training corpus for linguistic annotation of modern standard Slovene has been in continuous development for 15 years. |
| Approach: | They introduce an upgrade of a training corpus for linguistic annotation of modern standard Slovene. |
| Outcome: | The revised corpus, built on its predecessor, doubles in size and depth of annotation layers. |
DELTA: A Toolkit for Measuring Linguistic Diversity in Dependency-Parsed Corpora (2026.eacl-demo)
Copied to clipboard
| Challenge: | Existing tools for measuring diversity of specific linguistic phenomena are limited . we present an open-source framework for measuring linguistic diversity . |
| Approach: | They propose an open-source framework that integrates dependency tree querying with diversity computation. |
| Outcome: | The proposed framework can measure diversity across multiple linguistic levels and dimensions. |