Papers by Ondřej Herman

2 papers
ShadowSense: A Multi-annotated Dataset for Evaluating Word Sense Induction (2024.lrec-main)

Copied to clipboard

Challenge: Existing word sense induction datasets are annotated by multiple annotators whose inter-annotator agreement is key reliability score .
Approach: They propose a dataset that is annotated by multiple annotators with a key reliability score for evaluation of systems automatically inducing word senses.
Outcome: The proposed dataset shows that it is more reliable than existing paradigms for word sense induction evaluation.
Detecting Subtle Sense Shift with Polysemy-Aware Trends (2026.eacl-short)

Copied to clipboard

Challenge: Existing research on lexical semantic change focuses on century-scale, well-curated corpora and binary "changed / unchanged" judgements.
Approach: They propose a language-independent pipeline that detects word-sense shifts in large, time-stamped web corpora.
Outcome: The proposed pipeline detects word-sense shifts in large, time-stamped web corpora.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations