Papers by Philipp Heinrich
Corpus Query Lingua Franca part II: Ontology (2020.lrec-1)
Copied to clipboard
| Challenge: | outlines the projected second part of the Corpus Query Lingua Franca (CQLF) family of standards . the existence of a large number of different corpus query languages poses an epistemic challenge for the research community . |
| Approach: | They propose to standardize the Corpus Query Lingua Franca (CQLF) family of standards . they present the assumptions and aims of the CQLF Metamodel and its basic structure . |
| Outcome: | The proposed second part of the Corpus Query Lingua Franca (CQLF) family is in the process of standardization at the International Standards Organization (ISO) the first part of CQLF Ontology was adopted as an international standard at the beginning of 2018 . |
A Corpus of German Reddit Exchanges (GeRedE) (2020.lrec-1)
Copied to clipboard
| Challenge: | Reddit is a popular online platform combining social news aggregation, discussion and microblogging. |
| Approach: | They propose a method to filter out German data and further pre-processing steps to find out what is linguistically peculiar in the German data. |
| Outcome: | The proposed method filters out German data and includes metadata and annotation layers. |
EmpiriST Corpus 2.0: Adding Manual Normalization, Lemmatization and Semantic Tagging to a German Web and CMC Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |
| Approach: | They extend the corpus with manually tokenized annotation layers for word form normalization, lemmatization and lexical semantics. |
| Outcome: | The EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |
Automatic Identification of COVID-19-Related Conspiracy Narratives in German Telegram Channels and Chats (2024.lrec-main)
Copied to clipboard
Philipp Heinrich, Andreas Blombach, Bao Minh Doan Dang, Leonardo Zilio, Linda Havenstein, Nathan Dykes, Stephanie Evert, Fabian Schäfer
| Challenge: | Existing methods to identify and track conspiracy narratives are difficult to track and use because of their short-lived nature. |
| Approach: | They analysed 1,000 German Telegram posts tagged with 14 fine-grained conspiracy narrative labels by three independent annotators. |
| Outcome: | The proposed methods compare well with off-the-shelf methods and human performance. |