Papers by Luise Modersohn
Towards Label-Agnostic Emotion Embeddings (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing representation schemes for emotion analysis are based on label formats, natural languages, and even disparate model architectures. |
| Approach: | They propose a training scheme that learns a shared latent representation of emotion independent from different label formats, natural languages, and even disparate model architectures. |
| Outcome: | The proposed model performs well on a wide range of datasets without penalizing prediction quality. |
GGPONC 2.0 - The German Clinical Guideline Corpus for Oncology: Curation Workflow, Annotation Policy, Baseline NER Taggers (2022.lrec-1)
Copied to clipboard
Florian Borchert, Christina Lohr, Luise Modersohn, Jonas Witt, Thomas Langer, Markus Follmann, Matthias Gietzelt, Bert Arnrich, Udo Hahn, Matthieu-P. Schapranow
| Challenge: | despite advances in language resources, there is still a shortage of annotated corpora covering (German) medical language. |
| Approach: | They propose to build on clinical guidelines with an annotation scheme based on SNOMED CT . they also train named entity recognition models on the new data set . |
| Outcome: | The new corpus can be built upon clinical guidelines with reasonable coverage of medical terminology. |
GerMedIQ: A Resource for Simulated and Synthesized Anamnesis Interview Responses in German (2025.acl-srw)
Copied to clipboard
Justin Hofenbitzer, Sebastian Schöning, Belle Sebastian, Jacqueline Lammert, Luise Modersohn, Martin Boeker, Diego Frassinelli
| Challenge: | Text corpora in non-English clinical contexts is scarce due to privacy restrictions and restricted access to secure environments. |
| Approach: | They propose to use Large Language Models to generate synthetic data using a German medical interview questions corpus. |
| Outcome: | The proposed dataset generates comparable responses to human-generated questions. |
ProGene - A Large-scale, High-Quality Protein-Gene Annotated Benchmark Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Genes and proteins are fundamental entities of molecular genetics and are important for precision medicine. |
| Approach: | They propose to use a corpus of gene and protein names to cope with this class of named entities in a large-scale annotation campaign at the Jena University Language & Information Engineering lab. |
| Outcome: | The proposed corpus is an overall subdomain-independent corpus . it consists of 3,308 MEDLINE abstracts with over 36k sentences and more than 960k tokens annotated with nearly 60k named entity mentions. |