Papers by Slobodan Beliga
Croatian Idioms Integration: Enhancing the LIdioms Multilingual Linked Idioms Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing datasets that include idioms from English, German, Italian, Portuguese and Russian do not include a comprehensive representation of idiomatic expressions in Croatian. |
| Approach: | They propose to extend existing RDF-based multilingual representation of idioms to include 1,042 Croatian idiomes in an Ontolex Lemon format. |
| Outcome: | The proposed resource includes 1,042 Croatian idioms in an Ontolex Lemon format to foster translation initiatives and facilitate intercultural exchange. |
Evaluation of Croatian Word Embeddings (L18-1)
Copied to clipboard
| Challenge: | Currently, research is focusing mostly on English. |
| Approach: | They propose to use word analogy datasets to evaluate word similarities in Croatian . they use Word2Vec and FastText to create word analogies from highdimensional space . |
| Outcome: | The proposed datasets show that word embeddings are able to capture the syntactic and semantic relationship between words. |