Papers by Natalie Dykes
A Corpus of German Reddit Exchanges (GeRedE) (2020.lrec-1)
Copied to clipboard
| Challenge: | Reddit is a popular online platform combining social news aggregation, discussion and microblogging. |
| Approach: | They propose a method to filter out German data and further pre-processing steps to find out what is linguistically peculiar in the German data. |
| Outcome: | The proposed method filters out German data and includes metadata and annotation layers. |
EmpiriST Corpus 2.0: Adding Manual Normalization, Lemmatization and Semantic Tagging to a German Web and CMC Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |
| Approach: | They extend the corpus with manually tokenized annotation layers for word form normalization, lemmatization and lexical semantics. |
| Outcome: | The EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |