Papers by Stefan Evert
Corpus Query Lingua Franca part II: Ontology (2020.lrec-1)
Copied to clipboard
| Challenge: | outlines the projected second part of the Corpus Query Lingua Franca (CQLF) family of standards . the existence of a large number of different corpus query languages poses an epistemic challenge for the research community . |
| Approach: | They propose to standardize the Corpus Query Lingua Franca (CQLF) family of standards . they present the assumptions and aims of the CQLF Metamodel and its basic structure . |
| Outcome: | The proposed second part of the Corpus Query Lingua Franca (CQLF) family is in the process of standardization at the International Standards Organization (ISO) the first part of CQLF Ontology was adopted as an international standard at the beginning of 2018 . |
EmpiriST Corpus 2.0: Adding Manual Normalization, Lemmatization and Semantic Tagging to a German Web and CMC Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |
| Approach: | They extend the corpus with manually tokenized annotation layers for word form normalization, lemmatization and lexical semantics. |
| Outcome: | The EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |
Delta vs. N-Gram Tracing: Evaluating the Robustness of Authorship Attribution Methods (L18-1)
Copied to clipboard
| Challenge: | a novel authorship attribution method is developed for short texts . delta measures are well-established, but N-gram tracing is not robust enough . |
| Approach: | They propose to use delta measures and N-gram tracing to compare short texts . they find they are highly sensitive to the choice of authors and texts in the corpus . |
| Outcome: | The proposed methods are highly sensitive to the selection of authors and texts in the comparison corpus. |