Towards Building the LEMI Readability Platform for Children’s Literature in the Romanian Language (2024.lrec-main)
Copied to clipboard
Madalina Chitez, Mihai Dascalu, Aura Cristina Udrea, Cosmin Strilețchi, Karla Csürös, Roxana Rogobete, Alexandru Oravițan
| Challenge: | Currently, no existing platform integrates a research-based readability formula for the Romanian language, making this tool unique. |
| Approach: | They propose a new readability tool for children’s literature in the Romanian language that uses a self-compiled corpus and a text analysis interface to generate automatic readability reports for uploaded short texts. |
| Outcome: | The proposed readability tool is specifically targeted at primary school students aged 7-11 . it extracts, tests, and calibrates a readability formula for Romanian using the children’s literature corpus and the platform functionalities. |
Similar Papers
Comparing and Developing Tools to Measure the Readability of Domain-Specific Texts (D19-1)
Copied to clipboard
Elissa Redmiles, Lisa Maszkiewicz, Emily Hwang, Dhruv Kuchhal, Everest Liu, Miraida Morales, Denis Peskov, Sudha Rao, Rock Stevens, Kristina Gligorić, Sean Kross, Michelle Mazurek, Hal Daumé III
| Challenge: | Despite this, we lack a thorough understanding of how to validly measure readability at scale, especially for domain-specific texts. |
| Approach: | They present a comparison of the validity of well-known readability measures and introduce a novel approach to measure readability at scale. |
| Outcome: | The proposed approach addresses shortcomings of existing measures. |
Modeling the Readability of German Targeting Adults and Children: An empirically broad analysis and its cross-corpus validation (C18-1)
Copied to clipboard
| Challenge: | a new corpus of german news broadcast subtitles is compiled and crawled . readability assessment is a task of linking a text to the appropriate target audience based on its complexity. |
| Approach: | They analyze two German educational media texts targeting adults and children . they use 400 automatically extracted measures of linguistic complexity from a wide range of linguistic domains . their most successful binary classification model for german readability shows high accuracy . |
| Outcome: | The proposed model shows high accuracy between 89.4%–98.9% for both data sets. |
A Bird’s-eye View of Language Processing Projects at the Romanian Academy (L18-1)
Copied to clipboard
| Challenge: | a recent article outlines five projects that address contemporary Romanian language . the authors argue that a constant accumulation of human expertise is needed to develop complex projects. |
| Approach: | a new article gives a general overview of five AI language-related projects at the Romanian Academy . they focus on the creation of a contemporary Romanian language text and speech corpus and language related applications . |
| Outcome: | a new article gives an overview of five AI language-related projects at the Romanian Academy . the projects address contemporary Romanian language, as well as language related applications . |
Mama/Papa, Is this Text for Me? (2020.coling-main)
Copied to clipboard
Rashedur Rahman, Gwénolé Lecorvé, Aline Étienne, Delphine Battistelli, Nicolas Béchet, Jonathan Chevelu
| Challenge: | Existing methods to predict minimal age from which text can be understood for children are unresolved in computational linguistics. |
| Approach: | They propose a method which predicts the minimum age from which a text can be understood by a recurrent neural network. |
| Outcome: | The proposed method outperforms state-of-the-art models at sentence and text levels and achieves mean absolute errors of 1.86 and 2.28. |
Alector: A Parallel Corpus of Simplified French Texts with Alignments of Misreadings by Poor and Dyslexic Readers (2020.lrec-1)
Copied to clipboard
| Challenge: | Typical readers tend to progress quickly in reading because of the automatic process, which increases word identification and vice-versa. |
| Approach: | They propose a parallel corpus for reading tests and for the development of automatic text simplification tools for children with reading difficulties. |
| Outcome: | The proposed corpus is available for consultation through a web interface and available on demand for research purposes. |
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)
Copied to clipboard
| Challenge: | Using monolingual-only data, we can automate readability assessment and text simplification of simplified language. |
| Approach: | They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data. |
| Outcome: | The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images. |
FABRA: French Aggregator-Based Readability Assessment toolkit (2022.lrec-1)
Copied to clipboard
Rodrigo Wilkens, David Alfter, Xiaoou Wang, Alice Pintard, Anaïs Tack, Kevin P. Yancey, Thomas François
| Challenge: | a large number of readability predictor variables are used to predict reading difficulty of texts . the most important predictors for native texts are lexical diversity, dependency counts and text coherence . |
| Approach: | They propose a readability toolkit based on aggregation of readability predictor variables . they show which features are most predictive on two different corpora . |
| Outcome: | The proposed toolkit improves performance over standard feature-based readability prediction. |
Classic4Children: Adapting Chinese Literary Classics for Children with Large Language Model (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent large language models (LLMs) overlook children’s reading preferences, which poses challenges in CLA. |
| Approach: | They propose a method that augments large language models with children's reading preferences for adaptation by obtaining characters' personalities and narrative structure as additional information for fine-grained instruction tuning. |
| Outcome: | The proposed method significantly improves performance in automatic and human evaluation. |
Assessing French Readability for Adults with Low Literacy: A Global and Local Perspective (2025.emnlp-main)
Copied to clipboard
| Challenge: | illiterate individuals are persons aged 15 years and above who cannot read and write with understanding a short simple statement on their everyday life. |
| Approach: | They propose a novel approach to assess french text readability for adults with low literacy skills using a global and segment-level difficulty scale. |
| Outcome: | The proposed approach addresses both global (full-text) and local (segment-level) difficulty scales. |
KidLM: Advancing Language Models for Children – Early Insights and Future Directions (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models have been shown to be effective in creating educational tools for children, yet there are significant challenges in maintaining key child-specific properties such as linguistic nuances, cognitive needs, and safety standards. |
| Approach: | They propose a user-centric data collection pipeline that involves gathering and validating a corpus specifically written for and sometimes by children. |
| Outcome: | The proposed model excels in understanding lower grade-level text, maintains safety by avoiding stereotypes, and captures children’s unique preferences. |