Papers by Filip Klubička
Is it worth it? Budget-related evaluation metrics for model selection (L18-1)
Copied to clipboard
| Challenge: | linguistic resources can be labor-intensive, requiring great amounts of work-hours and expert annotation. |
| Approach: | They propose a machine learning model that pre-annotates or filters content before annotating it . they argue that the model with the highest F-score may not have best separation . |
| Outcome: | a case study shows that the model with the highest F-score does not yield the highest profits . the model that has the highest score does not produce the highest profit, the study shows . |
English WordNet Random Walk Pseudo-Corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | a random walk over the WordNet taxonomy generates a set of pseudo-corpora . a resource description paper describes the creation and properties of such pseudo-corporates . |
| Approach: | They propose to use random walk to generate a set of pseudo-corpora over the English WordNet taxonomy. |
| Outcome: | The proposed pseudo-corpora can be used to train taxonomic word embeddings . the proposed pseudo corpora are generated from a random walk over the English wordnet taxonomy . |