Papers by Jaka Čibej

4 papers
Gigafida 2.0: The Reference Corpus of Written Standard Slovene (2020.lrec-1)

Copied to clipboard

Challenge: Gigafida reference corpus of Slovene is updated with new material and tools . focus of upgrade was on transformation from general reference corp to standard reference corp .
Approach: We present a new version of the Gigafida reference corpus of Slovene . the upgrade includes new material and better tools for annotating it .
Outcome: The new version of the Gigafida reference corpus of Slovene is described . the whole Gigido corpus was deduplicated for the first time .
Creating Expert Knowledge by Relying on Language Learners: a Generic Approach for Mass-Producing Language Resources by Combining Implicit Crowdsourcing and Language Learning (2020.lrec-1)

Copied to clipboard

Challenge: Lack of wide-coverage and high-quality LRs is a longstanding issue in natural language processing (NLP) however, there are no large initiatives of similar scale for creating new LR or improving existing ones.
Approach: They propose a generic approach to combine implicit crowdsourcing and language learning to mass-produce language resources (LRs) they describe its core paradigm that consists in pairing specific types of LRs with specific exercises .
Outcome: The proposed approach can be used in several learning scenarios to produce a multitude of NLP resources and alleviate the long-standing issue of the lack of LRs.
SUK 1.0: A New Training Corpus for Linguistic Annotation of Modern Standard Slovene (2024.lrec-main)

Copied to clipboard

Challenge: a training corpus for linguistic annotation of modern standard Slovene has been in continuous development for 15 years.
Approach: They introduce an upgrade of a training corpus for linguistic annotation of modern standard Slovene.
Outcome: The revised corpus, built on its predecessor, doubles in size and depth of annotation layers.
SI-NLI: A Slovene Natural Language Inference Dataset and Its Evaluation (2024.lrec-main)

Copied to clipboard

Challenge: Existing datasets for natural language inference (NLI) are limited to English and a few other well-resourced languages.
Approach: They propose to use a dataset for natural language inference to extend the resources for the task.
Outcome: The proposed dataset is constructed from scratch using knowledgeable annotators with carefully crafted guidelines aiming to avoid common problems in existing datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations