Papers by Lukáš Svoboda

2 papers
Towards Universal Segmentations: UniSegments 1.0 (2022.lrec-1)

Copied to clipboard

Challenge: Existing data resources for morphological segmentation are limited to 32 languages . a large number of word forms exist, with some sub-parts being "recycled" many times .
Approach: They propose a multilingual data resource for morphological segmentation in 32 languages . they analyze diversity of how individual linguistic phenomena are captured across them .
Outcome: The proposed scheme is based on 17 existing data resources relevant for segmentation in 32 languages.
Evaluation of Croatian Word Embeddings (L18-1)

Copied to clipboard

Challenge: Currently, research is focusing mostly on English.
Approach: They propose to use word analogy datasets to evaluate word similarities in Croatian . they use Word2Vec and FastText to create word analogies from highdimensional space .
Outcome: The proposed datasets show that word embeddings are able to capture the syntactic and semantic relationship between words.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations