Papers by Rasmus Hvingelby
DaNE: A Named Entity Resource for Danish (2020.lrec-1)
Copied to clipboard
Rasmus Hvingelby, Amalie Brogaard Pauli, Maria Barrett, Christina Rosted, Lasse Malm Lidegaard, Anders Søgaard
| Challenge: | a named entity annotation for the Danish Universal Dependencies treebank is the largest publicly available named entity gold annotation. |
| Approach: | They propose a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme DaNE. |
| Outcome: | The proposed annotations improve Danish named entity recognition over a recent cross-lingual approach and over norwegian training set. |
Type B Reflexivization as an Unambiguous Testbed for Multilingual Multi-Task Gender Bias (2020.emnlp-main)
Copied to clipboard
| Challenge: | English challenge datasets highlight gender-ambiguous occurrences of ‘doctor’ as male doctors, but they are not useful for other languages. |
| Approach: | They propose to build multi-task challenge datasets for detecting gender bias that lead to unambiguously wrong model predictions for languages with type B reflexivization. |
| Outcome: | The proposed dataset can detect gender bias in languages with type B reflexivization and spans four languages and four NLP tasks. |
Towards a Gold Standard for Evaluating Danish Word Embeddings (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing word embedding models resemble semantic similarity solely by distribution, but there seems to be a need for future judgments to measure similarity in full context and along more than a single spectrum. |
| Approach: | They propose a model-agnostic similarity goal standard for evaluating Danish word embeddings based on human judgments made by 42 native speakers of Danish. |
| Outcome: | The goal standard is applied to evaluate Danish word embeddings on 42 native speakers of Danish. |