Papers by Christoph Purschke

3 papers
Capturing Regional Variation with Distributed Place Representations and Geographic Retrofitting (D18-1)

Copied to clipboard

Challenge: Dialects are one of the main drivers of language variation, a major challenge for natural language processing tools.
Approach: They use a corpus of 16.8M anonymous online posts to learn continuous document representations of cities.
Outcome: The proposed method matches dialect areas at different granularities against an existing dialect map.
Guided Distant Supervision for Multilingual Relation Extraction Data: Adapting to a New Language (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods to extract relationships are limited to English and require annotating datasets in order to be expensive and time-consuming.
Approach: They apply guided distant supervision to create a large biographical relationship extraction dataset for German using 80,000 instances for nine relationship types.
Outcome: The proposed dataset is the largest biographical German relationship extraction dataset.
ltzGLUE: Luxembourgish General Language Understanding Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: ltzGLUE is the first official NLU benchmark for Luxembourgish (LTZ) based on the popular GLUE benchmark for English.
Approach: They propose a new natural language understanding (NLU) benchmark for Luxembourgish based on the popular GLUE benchmark for English.
Outcome: The proposed model performs well across many languages and is based on the GLUE benchmark for English.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations