Papers by Christina Lohr
DOPA METER – A Tool Suite for Metrical Document Profiling and Aggregation (2023.emnlp-demo)
Copied to clipboard
| Challenge: | Xiao et al., 2022) examines the behavior of written language in a metrical way. |
| Approach: | They propose a tool suite for the metrical investigation of written language that provides diagnostic means for its division into discourse categories, such as registers, genres, and style. |
| Outcome: | The proposed system provides means for scoring linguistic behavior at the lexical, syntactic and semantic dimension. |
GGPONC 2.0 - The German Clinical Guideline Corpus for Oncology: Curation Workflow, Annotation Policy, Baseline NER Taggers (2022.lrec-1)
Copied to clipboard
Florian Borchert, Christina Lohr, Luise Modersohn, Jonas Witt, Thomas Langer, Markus Follmann, Matthias Gietzelt, Bert Arnrich, Udo Hahn, Matthieu-P. Schapranow
| Challenge: | despite advances in language resources, there is still a shortage of annotated corpora covering (German) medical language. |
| Approach: | They propose to build on clinical guidelines with an annotation scheme based on SNOMED CT . they also train named entity recognition models on the new data set . |
| Outcome: | The new corpus can be built upon clinical guidelines with reasonable coverage of medical terminology. |
Sharing Copies of Synthetic Clinical Corpora without Physical Distribution — A Case Study to Get Around IPRs and Privacy Constraints Featuring the German JSYNCC Corpus (L18-1)
Copied to clipboard
| Challenge: | eu legal culture imposes unsurmountable hurdles to exploit copyright protected language data . legal constraints have seriously hampered progress in resource-greedy NLP research . authors propose a new approach for the creation and re-use of clinical corpora . |
| Approach: | They propose a method for the creation and re-use of clinical corpora based on a two-step workflow . they substitute authentic clinical documents by synthetic ones, i.e., made-up reports and case studies . |
| Outcome: | a new approach replaces authentic clinical documents by synthetic ones, i.e., made-up reports and case studies published in medical e-textbooks. |
ProGene - A Large-scale, High-Quality Protein-Gene Annotated Benchmark Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Genes and proteins are fundamental entities of molecular genetics and are important for precision medicine. |
| Approach: | They propose to use a corpus of gene and protein names to cope with this class of named entities in a large-scale annotation campaign at the Jena University Language & Information Engineering lab. |
| Outcome: | The proposed corpus is an overall subdomain-independent corpus . it consists of 3,308 MEDLINE abstracts with over 36k sentences and more than 960k tokens annotated with nearly 60k named entity mentions. |