Papers with German-language

10 papers
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation (2026.eacl-long)

Copied to clipboard

Challenge: Recent studies show that data quality can significantly boost performance and training efficiency for large language models.
Approach: They propose a German-language dataset curation pipeline that combines heuristic and model-based filtering techniques with synthetic data generation.
Outcome: The proposed pipeline can be used to create a large-scale German pre-training dataset using common Crawl web data, fineweb2 and synthetically generated data conditioned on real, organic web data.
Extractive Summarisation for German-language Data: A Text-level Approach with Discourse Features (2022.coling-1)

Copied to clipboard

Challenge: Using RST, extractive summarisation involves using select phrases and sentences as a summary, which still remains a strong method for producing summaries despite its simple nature.
Approach: They propose to use RST-based features to analyse the connection between summary sentences and several RST features and transfer these insights to various automated summarisation models.
Outcome: The proposed models are based on the best features proposed over the last 20+ years and incorporate the best ones into the proposed models.
“Beste Grüße, Maria Meyer” — Pseudonymization of Privacy-Sensitive Information in Emails (2022.lrec-1)

Copied to clipboard

Challenge: exploding amount of user-generated content has spurred research to deal with documents from various digital communication formats.
Approach: They propose to identify text spans that carry information revealing an individual’s identity and substitute them with synthetically generated surrogates.
Outcome: The proposed model is based on a German-language email corpus and evaluates its training data on pseudonymized data.
Allgemeine Musikalische Zeitung as a Searchable Online Corpus (2020.lrec-1)

Copied to clipboard

Challenge: specialized newspapers are not well curated in terms of digitization quality, data formatting, completeness, redundancy (de-duplication), supply of metadata, and hence, searchability.
Approach: They propose a workflow that copes with a posteriori digitization problems, inconsistent OCRing and index building for searchability for a major German-language newspaper of the Romantic Age.
Outcome: The proposed workflow copes with a posteriori digitization problems, inconsistent OCRing and index building for searchability.
Acquiring a Formality-Informed Lexical Resource for Style Analysis (2021.eacl-main)

Copied to clipboard

Challenge: lexico-statistics analysis of formality levels in written communication has long been dominated by application concerns, such as authorship and plagiarism assignment problems.
Approach: They propose a lexicon with entries ordered by their degree of (in)formality and let crowdworkers assess the enlarged set of lexical items on a continuous informal-formal scale as a gold standard for evaluation.
Outcome: The proposed lexicon is evaluated on a German-language email corpus and is then evaluated by crowdworkers.
Detecting Scenes in Fiction: A new Segmentation Task (2021.eacl-main)

Copied to clipboard

Challenge: Text segmentation is a long standing issue in the area of natural language processing . even modern methods struggle with processing text longer than a couple of sentences or paragraphs .
Approach: They introduce the task of scene segmentation on narrative texts and provide an annotated corpus . they discuss linguistic and narrative properties of the task and provide baseline experiments .
Outcome: The proposed task is very challenging and the results are impressive.
A Corpus of German Citizen Contributions in Mobility Planning: Supporting Evaluation Through Multidimensional Classification (2022.lrec-1)

Copied to clipboard

Challenge: Political authorities in democratic countries consult the public in order to allow citizens to voice their ideas and concerns on specific issues.
Approach: They propose a publicly-available corpus that includes citizen contributions from six mobility-related planning processes in five german municipalities.
Outcome: The proposed corpus includes several thousand citizen contributions from six mobility-related planning processes in five German municipalities.
How Much Data Do You Need? About the Creation of a Ground Truth for Black Letter and the Effectiveness of Neural OCR (2020.lrec-1)

Copied to clipboard

Challenge: Recent advances in Optical Character Recognition and Handwritten Text Recognition have led to more accurate text recognition of historical documents.
Approach: They propose to build a ground truth for a German-language newspaper published in black letter . they also evaluate the performance of different OCR engines and estimate how much data is needed to achieve high-quality OCR results.
Outcome: The proposed model can recognise black letter text and performs well on data they have not seen during training.
taz2024full: Analysing German Newspapers for Gender Bias and Discrimination across Decades (2025.findings-acl)

Copied to clipboard

Challenge: a large corpus of German newspaper articles is available for free in other languages, such as English.
Approach: They propose to use taz2024full to analyse gender representation across four decades of reporting.
Outcome: The proposed corpus supports a wide range of applications from diachronic language analysis to critical media studies.
Is Gender Reference Gender-specific? Studies in a Polar Domain (2024.lrec-main)

Copied to clipboard

Challenge: a german-language newspaper corpus contains a large number of newspaper texts with gender tags . a gender-specific way of gender reference is investigated by using a polar load-based classifier .
Approach: They investigate how gender authorship influences positive and negative gender reference . they use a german valence lexicon, a German polar load lexicone and a verb-based analysis of the polar role a noun plays .
Outcome: The proposed method mainly uses a German valence lexicon, a german polar load lexical, and a polar lexicogramma.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations