Papers by Stefan Heinrich
Corpus Query Lingua Franca part II: Ontology (2020.lrec-1)
Copied to clipboard
| Challenge: | outlines the projected second part of the Corpus Query Lingua Franca (CQLF) family of standards . the existence of a large number of different corpus query languages poses an epistemic challenge for the research community . |
| Approach: | They propose to standardize the Corpus Query Lingua Franca (CQLF) family of standards . they present the assumptions and aims of the CQLF Metamodel and its basic structure . |
| Outcome: | The proposed second part of the Corpus Query Lingua Franca (CQLF) family is in the process of standardization at the International Standards Organization (ISO) the first part of CQLF Ontology was adopted as an international standard at the beginning of 2018 . |
EmpiriST Corpus 2.0: Adding Manual Normalization, Lemmatization and Semantic Tagging to a German Web and CMC Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |
| Approach: | They extend the corpus with manually tokenized annotation layers for word form normalization, lemmatization and lexical semantics. |
| Outcome: | The EmpiriST corpus is a manually tokenized and part-of-speech tagged corpus of German web and CMC data. |
Research Community Perspectives on “Intelligence” and Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Despite the widespread use of ‘artificial intelligence’ (AI) framing in NLP research, it is not clear what researchers mean by ”intelligence”. |
| Approach: | They propose to use the term "AI" to describe the perception of a system as intelligent, but note that it is not accepted by the majority of respondents. |
| Outcome: | The results suggest that the perception of the current NLP systems as 'intelligent' is a minority position (29%). |