Papers by Simon Krek

7 papers
Gigafida 2.0: The Reference Corpus of Written Standard Slovene (2020.lrec-1)

Copied to clipboard

Challenge: Gigafida reference corpus of Slovene is updated with new material and tools . focus of upgrade was on transformation from general reference corp to standard reference corp .
Approach: We present a new version of the Gigafida reference corpus of Slovene . the upgrade includes new material and better tools for annotating it .
Outcome: The new version of the Gigafida reference corpus of Slovene is described . the whole Gigido corpus was deduplicated for the first time .
Introducing the CURLICAT Corpora: Seven-language Domain Specific Annotated Corpora from Curated Sources (2022.lrec-1)

Copied to clipboard

Challenge: The CURLICAT CEF Telecom project aims to collect and deeply annotate a set of large corpora from selected domains.
Approach: They present the results of the CURLICAT CEF Telecom project . they propose to collect and deeply annotate a set of large corpora from selected domains .
Outcome: The CURLICAT CEF Telecom project provides a set of large corpora from selected domains . the corporatized corporates are tokenized, lemmatized and morphologically analysed .
What’s the Meaning of Superhuman Performance in Today’s NLU? (2023.acl-long)

Copied to clipboard

Challenge: Recent research has focused on developing larger pretrained language models and introducing benchmarks such as SuperGLUE and SQuAD to measure their abilities.
Approach: They propose to use benchmarks such as SuperGLUE and SQUAD to evaluate PLMs' abilities in language understanding, reasoning, and reading comprehension to assess their performance.
Outcome: The proposed benchmarks have serious limitations affecting comparison between humans and PLMs and provide recommendations for fairer and more transparent benchmarks.
A Multilingual Evaluation Dataset for Monolingual Word Sense Alignment (2020.lrec-1)

Copied to clipboard

Challenge: a new dataset aims to align monolingual dictionaries with a single sense level for 15 languages . this dataset covers a wide range of languages and resources .
Approach: They propose to manually align monolingual dictionaries with possible semantic relationships . they use 15 languages to create a new baseline for the task of monolingual word sense alignment .
Outcome: The proposed dataset covers 15 languages and covers the more challenging task of linking general-purpose language.
SUK 1.0: A New Training Corpus for Linguistic Annotation of Modern Standard Slovene (2024.lrec-main)

Copied to clipboard

Challenge: a training corpus for linguistic annotation of modern standard Slovene has been in continuous development for 15 years.
Approach: They introduce an upgrade of a training corpus for linguistic annotation of modern standard Slovene.
Outcome: The revised corpus, built on its predecessor, doubles in size and depth of annotation layers.
The MARCELL Legislative Corpus (2020.lrec-1)

Copied to clipboard

Challenge: MARCELL corpus provides a rich and valuable source for further studies and developments in machine learning, cross-lingual terminological data extraction and classification.
Approach: They present the results of the project MARCELL CEF Telecom . they aim to collect and deeply annotate a large comparable corpus of legal documents .
Outcome: The MARCELL corpus includes 7 monolingual sub-corpora containing the body of respective national legislative documents.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations