Papers by Luise Modersohn

4 papers
Towards Label-Agnostic Emotion Embeddings (2021.emnlp-main)

Copied to clipboard

Challenge: Existing representation schemes for emotion analysis are based on label formats, natural languages, and even disparate model architectures.
Approach: They propose a training scheme that learns a shared latent representation of emotion independent from different label formats, natural languages, and even disparate model architectures.
Outcome: The proposed model performs well on a wide range of datasets without penalizing prediction quality.
GGPONC 2.0 - The German Clinical Guideline Corpus for Oncology: Curation Workflow, Annotation Policy, Baseline NER Taggers (2022.lrec-1)

Copied to clipboard

Challenge: despite advances in language resources, there is still a shortage of annotated corpora covering (German) medical language.
Approach: They propose to build on clinical guidelines with an annotation scheme based on SNOMED CT . they also train named entity recognition models on the new data set .
Outcome: The new corpus can be built upon clinical guidelines with reasonable coverage of medical terminology.
GerMedIQ: A Resource for Simulated and Synthesized Anamnesis Interview Responses in German (2025.acl-srw)

Copied to clipboard

Challenge: Text corpora in non-English clinical contexts is scarce due to privacy restrictions and restricted access to secure environments.
Approach: They propose to use Large Language Models to generate synthetic data using a German medical interview questions corpus.
Outcome: The proposed dataset generates comparable responses to human-generated questions.
ProGene - A Large-scale, High-Quality Protein-Gene Annotated Benchmark Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Genes and proteins are fundamental entities of molecular genetics and are important for precision medicine.
Approach: They propose to use a corpus of gene and protein names to cope with this class of named entities in a large-scale annotation campaign at the Jena University Language & Information Engineering lab.
Outcome: The proposed corpus is an overall subdomain-independent corpus . it consists of 3,308 MEDLINE abstracts with over 36k sentences and more than 960k tokens annotated with nearly 60k named entity mentions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations