Papers with German-language
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation (2026.eacl-long)
Copied to clipboard
Thomas F Burns, Letitia Parcalabescu, Stephan Waeldchen, Michael Barlow, Gregor Ziegltrum, Volker Stampa, Bastian Harren, Björn Deiseroth
| Challenge: | Recent studies show that data quality can significantly boost performance and training efficiency for large language models. |
| Approach: | They propose a German-language dataset curation pipeline that combines heuristic and model-based filtering techniques with synthetic data generation. |
| Outcome: | The proposed pipeline can be used to create a large-scale German pre-training dataset using common Crawl web data, fineweb2 and synthetically generated data conditioned on real, organic web data. |
Extractive Summarisation for German-language Data: A Text-level Approach with Discourse Features (2022.coling-1)
Copied to clipboard
| Challenge: | Using RST, extractive summarisation involves using select phrases and sentences as a summary, which still remains a strong method for producing summaries despite its simple nature. |
| Approach: | They propose to use RST-based features to analyse the connection between summary sentences and several RST features and transfer these insights to various automated summarisation models. |
| Outcome: | The proposed models are based on the best features proposed over the last 20+ years and incorporate the best ones into the proposed models. |
“Beste Grüße, Maria Meyer” — Pseudonymization of Privacy-Sensitive Information in Emails (2022.lrec-1)
Copied to clipboard
| Challenge: | exploding amount of user-generated content has spurred research to deal with documents from various digital communication formats. |
| Approach: | They propose to identify text spans that carry information revealing an individual’s identity and substitute them with synthetically generated surrogates. |
| Outcome: | The proposed model is based on a German-language email corpus and evaluates its training data on pseudonymized data. |
Allgemeine Musikalische Zeitung as a Searchable Online Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | specialized newspapers are not well curated in terms of digitization quality, data formatting, completeness, redundancy (de-duplication), supply of metadata, and hence, searchability. |
| Approach: | They propose a workflow that copes with a posteriori digitization problems, inconsistent OCRing and index building for searchability for a major German-language newspaper of the Romantic Age. |
| Outcome: | The proposed workflow copes with a posteriori digitization problems, inconsistent OCRing and index building for searchability. |
Acquiring a Formality-Informed Lexical Resource for Style Analysis (2021.eacl-main)
Copied to clipboard
| Challenge: | lexico-statistics analysis of formality levels in written communication has long been dominated by application concerns, such as authorship and plagiarism assignment problems. |
| Approach: | They propose a lexicon with entries ordered by their degree of (in)formality and let crowdworkers assess the enlarged set of lexical items on a continuous informal-formal scale as a gold standard for evaluation. |
| Outcome: | The proposed lexicon is evaluated on a German-language email corpus and is then evaluated by crowdworkers. |
Detecting Scenes in Fiction: A new Segmentation Task (2021.eacl-main)
Copied to clipboard
Albin Zehe, Leonard Konle, Lea Katharina Dümpelmann, Evelyn Gius, Andreas Hotho, Fotis Jannidis, Lucas Kaufmann, Markus Krug, Frank Puppe, Nils Reiter, Annekea Schreiber, Nathalie Wiedmer
| Challenge: | Text segmentation is a long standing issue in the area of natural language processing . even modern methods struggle with processing text longer than a couple of sentences or paragraphs . |
| Approach: | They introduce the task of scene segmentation on narrative texts and provide an annotated corpus . they discuss linguistic and narrative properties of the task and provide baseline experiments . |
| Outcome: | The proposed task is very challenging and the results are impressive. |
A Corpus of German Citizen Contributions in Mobility Planning: Supporting Evaluation Through Multidimensional Classification (2022.lrec-1)
Copied to clipboard
| Challenge: | Political authorities in democratic countries consult the public in order to allow citizens to voice their ideas and concerns on specific issues. |
| Approach: | They propose a publicly-available corpus that includes citizen contributions from six mobility-related planning processes in five german municipalities. |
| Outcome: | The proposed corpus includes several thousand citizen contributions from six mobility-related planning processes in five German municipalities. |
How Much Data Do You Need? About the Creation of a Ground Truth for Black Letter and the Effectiveness of Neural OCR (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent advances in Optical Character Recognition and Handwritten Text Recognition have led to more accurate text recognition of historical documents. |
| Approach: | They propose to build a ground truth for a German-language newspaper published in black letter . they also evaluate the performance of different OCR engines and estimate how much data is needed to achieve high-quality OCR results. |
| Outcome: | The proposed model can recognise black letter text and performs well on data they have not seen during training. |
taz2024full: Analysing German Newspapers for Gender Bias and Discrimination across Decades (2025.findings-acl)
Copied to clipboard
| Challenge: | a large corpus of German newspaper articles is available for free in other languages, such as English. |
| Approach: | They propose to use taz2024full to analyse gender representation across four decades of reporting. |
| Outcome: | The proposed corpus supports a wide range of applications from diachronic language analysis to critical media studies. |
Is Gender Reference Gender-specific? Studies in a Polar Domain (2024.lrec-main)
Copied to clipboard
| Challenge: | a german-language newspaper corpus contains a large number of newspaper texts with gender tags . a gender-specific way of gender reference is investigated by using a polar load-based classifier . |
| Approach: | They investigate how gender authorship influences positive and negative gender reference . they use a german valence lexicon, a German polar load lexicone and a verb-based analysis of the polar role a noun plays . |
| Outcome: | The proposed method mainly uses a German valence lexicon, a german polar load lexical, and a polar lexicogramma. |