Papers by Ondřej Herman
ShadowSense: A Multi-annotated Dataset for Evaluating Word Sense Induction (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing word sense induction datasets are annotated by multiple annotators whose inter-annotator agreement is key reliability score . |
| Approach: | They propose a dataset that is annotated by multiple annotators with a key reliability score for evaluation of systems automatically inducing word senses. |
| Outcome: | The proposed dataset shows that it is more reliable than existing paradigms for word sense induction evaluation. |
Detecting Subtle Sense Shift with Polysemy-Aware Trends (2026.eacl-short)
Copied to clipboard
| Challenge: | Existing research on lexical semantic change focuses on century-scale, well-curated corpora and binary "changed / unchanged" judgements. |
| Approach: | They propose a language-independent pipeline that detects word-sense shifts in large, time-stamped web corpora. |
| Outcome: | The proposed pipeline detects word-sense shifts in large, time-stamped web corpora. |