Subjective Text Complexity Assessment for German (2022.lrec-1)

Copied to clipboard

Challenge: Often, readability is defined as how easily a written text is to read.
Approach: They propose to use a corpus of sentences provided by a German IT service provider to assess the readability of German text.
Outcome: The proposed model can predict complexity of German text by using linguistically motivated features.

Similar Papers

A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)

Copied to clipboard

Challenge: Using monolingual-only data, we can automate readability assessment and text simplification of simplified language.
Approach: They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data.
Outcome: The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images.
Modeling the Readability of German Targeting Adults and Children: An empirically broad analysis and its cross-corpus validation (C18-1)

Copied to clipboard

Challenge: a new corpus of german news broadcast subtitles is compiled and crawled . readability assessment is a task of linking a text to the appropriate target audience based on its complexity.
Approach: They analyze two German educational media texts targeting adults and children . they use 400 automatically extracted measures of linguistic complexity from a wide range of linguistic domains . their most successful binary classification model for german readability shows high accuracy .
Outcome: The proposed model shows high accuracy between 89.4%–98.9% for both data sets.
Evaluating Readability Metrics for German Medical Text Simplification (2025.coling-main)

Copied to clipboard

Challenge: Clinical reports and scientific information sources are written for medical experts, preventing patients from understanding the main messages of these texts.
Approach: They evaluated the suitability of 18 statistical, part-of-speech-based, syntactic, semantic and fluency metrics to measure readability of German medical texts.
Outcome: The proposed measures are compared with standard methods on English medical texts and simplified summaries.
Is this Sentence Difficult? Do you Agree? (D18-1)

Copied to clipboard

Challenge: a crowdsourcing-based approach to model sentence complexity is proposed . word-level predictors shown to correlate with greater processing difficulties are e.g. word frequency, age of acquisition, root frequency effect, orthographic neighbourhood frequency .
Approach: They propose a crowdsourcing-based approach to model human perception of sentence complexity using a corpus of sentences rated with judgments of complexity for two typologically-different languages.
Outcome: The proposed model predicts agreement among annotators independently from the assigned judgment and the perception of sentence complexity in Italian and English.
Estimating Lexical Complexity from Document-Level Distributions (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for complexity estimation are limited to entire documents . health assessment tools are too short for existing methods to apply .
Approach: They propose a two-step approach for estimating lexical complexity that does not rely on pre-annotated data.
Outcome: The proposed method is tested on the Norwegian language and compares with other assessment tools.
DETECT: Determining Ease and Textual Clarity of German Text Simplifications (2026.eacl-long)

Copied to clipboard

Challenge: Current evaluation of German automatic text simplification relies on general-purpose metrics such as SARI, BLEU, and BERTScore.
Approach: They propose a German-specific metric that holistically evaluates ATS quality across all three dimensions of simplicity, meaning preservation, and fluency.
Outcome: The proposed metric achieves higher correlations with human judgments than widely used ATS metrics.
Automatic Assessment of Conceptual Text Complexity Using Knowledge Graphs (C18-1)

Copied to clipboard

Challenge: Existing methods to assess text complexity only at lexical and syntactic levels have not been attempted.
Approach: They propose to automatically estimate conceptual complexity using graph-based measures on a large knowledge base.
Outcome: The proposed measures achieve high discriminative power even in a default setup.
A linguistically-motivated evaluation methodology for unraveling model’s abilities in reading comprehension tasks (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models fail for linguistic characteristics of input examples, despite the impressive quantity of scientific studies dedicated to them, the capabilities, limitations, and risks of these models remain largely unknown.
Approach: They propose to use semantic frame annotation to characterize examples by a small number of complexity factors to account for model’s difficulty.
Outcome: The proposed evaluation methodology is based on the intuition that certain examples consistently yield lower scores regardless of model size or architecture.
Using Eye-tracking Data to Predict the Readability of Brazilian Portuguese Sentences in Single-task, Multi-task and Sequential Transfer Learning Approaches (2020.coling-main)

Copied to clipboard

Challenge: Sentence complexity assessment is a relatively new task in Natural Language Processing.
Approach: They propose to use Brazilian Portuguese to evaluate sentences with linguistic features to improve readability.
Outcome: The proposed model reaches the state-of-the-art for Brazilian Portuguese with 97.8% accuracy with linguistic features.
CoCo: A Tool for Automatically Assessing Conceptual Complexity of Texts (2020.lrec-1)

Copied to clipboard

Challenge: Traditional text complexity assessment only takes into account lexical and lexiconal complexity.
Approach: They propose a tool for automatic assessment of conceptual text complexity based on the current state-of-the-art unsupervised approach . they compare the current implementation with the state of the art and discuss the influence of the choice of entity linker on the performance of the tool.
Outcome: The proposed tool can be personalized and adapted to the needs of struggling readers.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations