Papers by Kaja Dobrovoljc

5 papers
Gigafida 2.0: The Reference Corpus of Written Standard Slovene (2020.lrec-1)

Copied to clipboard

Challenge: Gigafida reference corpus of Slovene is updated with new material and tools . focus of upgrade was on transformation from general reference corp to standard reference corp .
Approach: We present a new version of the Gigafida reference corpus of Slovene . the upgrade includes new material and better tools for annotating it .
Outcome: The new version of the Gigafida reference corpus of Slovene is described . the whole Gigido corpus was deduplicated for the first time .
Spoken Language Treebanks in Universal Dependencies: an Overview (2022.lrec-1)

Copied to clipboard

Challenge: spoken language treebanks have divergent annotation schemes limiting cross-resource explorations . many spoken language trees have no written form, but many of the world languages have no spoken form at all.
Approach: They propose to use the Universal Dependencies annotation scheme to annotate spoken language treebanks using a morphosyntactic annotation scheme.
Outcome: The proposed treebanks differ significantly with respect to the inventory and format of transcribed phenomena and the principles adopted in their morphosyntactic annotation.
Gos 2: A New Reference Corpus of Spoken Slovenian (2024.lrec-main)

Copied to clipboard

Challenge: a new corpus of spoken Slovenian has been added to the Gos reference corpus . the corpus is now more than double the original size of 300 hours, 2.4 million words .
Approach: They propose to add speech recordings and transcriptions from two related initiatives, the Gos VideoLectures corpus of public academic speech, and the Artur speech recognition database.
Outcome: The new corpus is double the original size and contains 2.4 million words . it includes speech recordings and transcriptions from two related initiatives .
SUK 1.0: A New Training Corpus for Linguistic Annotation of Modern Standard Slovene (2024.lrec-main)

Copied to clipboard

Challenge: a training corpus for linguistic annotation of modern standard Slovene has been in continuous development for 15 years.
Approach: They introduce an upgrade of a training corpus for linguistic annotation of modern standard Slovene.
Outcome: The revised corpus, built on its predecessor, doubles in size and depth of annotation layers.
DELTA: A Toolkit for Measuring Linguistic Diversity in Dependency-Parsed Corpora (2026.eacl-demo)

Copied to clipboard

Challenge: Existing tools for measuring diversity of specific linguistic phenomena are limited . we present an open-source framework for measuring linguistic diversity .
Approach: They propose an open-source framework that integrates dependency tree querying with diversity computation.
Outcome: The proposed framework can measure diversity across multiple linguistic levels and dimensions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations