Papers by Christina Lohr

4 papers
DOPA METER – A Tool Suite for Metrical Document Profiling and Aggregation (2023.emnlp-demo)

Copied to clipboard

Challenge: Xiao et al., 2022) examines the behavior of written language in a metrical way.
Approach: They propose a tool suite for the metrical investigation of written language that provides diagnostic means for its division into discourse categories, such as registers, genres, and style.
Outcome: The proposed system provides means for scoring linguistic behavior at the lexical, syntactic and semantic dimension.
GGPONC 2.0 - The German Clinical Guideline Corpus for Oncology: Curation Workflow, Annotation Policy, Baseline NER Taggers (2022.lrec-1)

Copied to clipboard

Challenge: despite advances in language resources, there is still a shortage of annotated corpora covering (German) medical language.
Approach: They propose to build on clinical guidelines with an annotation scheme based on SNOMED CT . they also train named entity recognition models on the new data set .
Outcome: The new corpus can be built upon clinical guidelines with reasonable coverage of medical terminology.
Sharing Copies of Synthetic Clinical Corpora without Physical Distribution — A Case Study to Get Around IPRs and Privacy Constraints Featuring the German JSYNCC Corpus (L18-1)

Copied to clipboard

Challenge: eu legal culture imposes unsurmountable hurdles to exploit copyright protected language data . legal constraints have seriously hampered progress in resource-greedy NLP research . authors propose a new approach for the creation and re-use of clinical corpora .
Approach: They propose a method for the creation and re-use of clinical corpora based on a two-step workflow . they substitute authentic clinical documents by synthetic ones, i.e., made-up reports and case studies .
Outcome: a new approach replaces authentic clinical documents by synthetic ones, i.e., made-up reports and case studies published in medical e-textbooks.
ProGene - A Large-scale, High-Quality Protein-Gene Annotated Benchmark Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Genes and proteins are fundamental entities of molecular genetics and are important for precision medicine.
Approach: They propose to use a corpus of gene and protein names to cope with this class of named entities in a large-scale annotation campaign at the Jena University Language & Information Engineering lab.
Outcome: The proposed corpus is an overall subdomain-independent corpus . it consists of 3,308 MEDLINE abstracts with over 36k sentences and more than 960k tokens annotated with nearly 60k named entity mentions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations