Papers by Christoph Purschke
Capturing Regional Variation with Distributed Place Representations and Geographic Retrofitting (D18-1)
Copied to clipboard
| Challenge: | Dialects are one of the main drivers of language variation, a major challenge for natural language processing tools. |
| Approach: | They use a corpus of 16.8M anonymous online posts to learn continuous document representations of cities. |
| Outcome: | The proposed method matches dialect areas at different granularities against an existing dialect map. |
Guided Distant Supervision for Multilingual Relation Extraction Data: Adapting to a New Language (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods to extract relationships are limited to English and require annotating datasets in order to be expensive and time-consuming. |
| Approach: | They apply guided distant supervision to create a large biographical relationship extraction dataset for German using 80,000 instances for nine relationship types. |
| Outcome: | The proposed dataset is the largest biographical German relationship extraction dataset. |
ltzGLUE: Luxembourgish General Language Understanding Evaluation (2026.findings-acl)
Copied to clipboard
Alistair Plum, Felicia Körner, Anne-Marie Lutgen, Laura Bernardy, Fred Philippy, Emilia Milano, Nils Rehlinger, Cedric Lothritz, Tharindu Ranasinghe, Barbara Plank, Christoph Purschke
| Challenge: | ltzGLUE is the first official NLU benchmark for Luxembourgish (LTZ) based on the popular GLUE benchmark for English. |
| Approach: | They propose a new natural language understanding (NLU) benchmark for Luxembourgish based on the popular GLUE benchmark for English. |
| Outcome: | The proposed model performs well across many languages and is based on the GLUE benchmark for English. |