Challenge: Using monolingual-only data, we can automate readability assessment and text simplification of simplified language.
Approach: They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data.
Outcome: The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images.

Similar Papers

Modeling the Readability of German Targeting Adults and Children: An empirically broad analysis and its cross-corpus validation (C18-1)

Copied to clipboard

Challenge: a new corpus of german news broadcast subtitles is compiled and crawled . readability assessment is a task of linking a text to the appropriate target audience based on its complexity.
Approach: They analyze two German educational media texts targeting adults and children . they use 400 automatically extracted measures of linguistic complexity from a wide range of linguistic domains . their most successful binary classification model for german readability shows high accuracy .
Outcome: The proposed model shows high accuracy between 89.4%–98.9% for both data sets.
Language Models for German Text Simplification: Overcoming Parallel Data Scarcity through Style-specific Pre-training (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to train automatic text simplification systems for languages other than English are limited by the lack of parallel data.
Approach: They propose to use German Easy Language as a corpus of automatic text simplification systems to fine-tune language models to the style characteristics of the language.
Outcome: The proposed language models adapt to the style characteristics of Easy Language and output more accessible texts.
DEplain: A German Parallel Corpus with Intralingual Translations into Plain Language for Sentence and Document Simplification (2023.acl-long)

Copied to clipboard

Challenge: Current text simplification research mostly focuses on English and on sentencelevel simplification.
Approach: They propose to use a dataset of parallel, professionally written and manually aligned simplifications in plain German "plain DE" and "Einfache Sprache" they build a web harvester and experiment with automatic alignment methods to facilitate integration of non-aligned and to be-published parallel documents.
Outcome: The proposed dataset of parallel, professionally written and manually aligned simplifications in plain German is extended to 750 document pairs and 3.5k sentence pairs.
The SAMER Arabic Text Simplification Corpus (2024.lrec-main)

Copied to clipboard

Challenge: Our corpus includes 159K words selected from 15 publicly available Arabic fiction novels . text simplification aims to reduce the complexity of a text while maintaining the overall grammaticality and core content.
Approach: They propose to annotate Arabic parallel corpus for text simplification targeting school-aged learners.
Outcome: The SAMER Corpus includes readability level annotations at both the document and word levels, as well as two simplified parallel versions for each text targeting learners at two different readability levels.
DETECT: Determining Ease and Textual Clarity of German Text Simplifications (2026.eacl-long)

Copied to clipboard

Challenge: Current evaluation of German automatic text simplification relies on general-purpose metrics such as SARI, BLEU, and BERTScore.
Approach: They propose a German-specific metric that holistically evaluates ATS quality across all three dimensions of simplicity, meaning preservation, and fluency.
Outcome: The proposed metric achieves higher correlations with human judgments than widely used ATS metrics.
Text Simplification from Professionally Produced Corpora (L18-1)

Copied to clipboard

Challenge: Existing approaches to Text Simplification rely on the Wikipedia-Simple Wikipedia parallel corpus, which is used for many tasks.
Approach: They propose to use the Newsela corpus to extract 550, 644 complex-simple sentence pairs from the corpus and introduce a lexical simplifier that uses the corpu to generate candidate simplifications.
Outcome: The proposed model outperforms state-of-the-art approaches and generates candidate simplifications from the newsela corpus.
Evaluating Readability Metrics for German Medical Text Simplification (2025.coling-main)

Copied to clipboard

Challenge: Clinical reports and scientific information sources are written for medical experts, preventing patients from understanding the main messages of these texts.
Approach: They evaluated the suitability of 18 statistical, part-of-speech-based, syntactic, semantic and fluency metrics to measure readability of German medical texts.
Outcome: The proposed measures are compared with standard methods on English medical texts and simplified summaries.
A Nontrivial Sentence Corpus for the Task of Sentence Readability Assessment in Portuguese (C18-1)

Copied to clipboard

Challenge: Effective textual communication depends on readers being proficient enough to comprehend texts . when meaning is not well conveyed, many losses and damages may occur .
Approach: They propose automatic evaluation of sentence readability task in Portuguese to improve readability.
Outcome: The proposed method correctly identifies the ranking of sentence pairs with an accuracy of 74.2%.
Klexikon: A German Dataset for Joint Summarization and Simplification (2022.lrec-1)

Copied to clipboard

Challenge: Traditionally, Text Simplification is a monolingual translation task where individual sentences are "translated" into a simplified version.
Approach: They propose to use a dataset to jointly simplify long source documents by combining sentences from a source and their simplified counterparts.
Outcome: The proposed system can summarize and simplify long source documents using almost 2,900 documents.
A Corpus of German Abstract Meaning Representation (DeAMR) (2024.lrec-main)

Copied to clipboard

Challenge: Abstract Meaning Representations (AMRs) are semantic graphs that abstract away from surface syntax and capture the meaning of who does what to whom in a sentence.
Approach: They propose to use German Abstract Meaning Representation (Deutsche AMR) to represent the structure and semantics of German.
Outcome: The proposed framework is based on an annotated corpus of 400 DeAMR in German and is validated through inter-annotator agreement.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations