Challenge: a study on diagnosing the textual difficulties of children's books is published . it focuses on the passages of the books that are difficult to understand for underage children .
Approach: They propose to diagnose the difficulties appearing in French children's books . they focus on the subject pronouns "il" and "elle" and detect difficult anaphoras .
Outcome: The proposed method detects half of the difficult anaphorical pronouns in french children's books . authors say it is complementary of previous approaches to support dyslexia .

Similar Papers

Alector: A Parallel Corpus of Simplified French Texts with Alignments of Misreadings by Poor and Dyslexic Readers (2020.lrec-1)

Copied to clipboard

Challenge: Typical readers tend to progress quickly in reading because of the automatic process, which increases word identification and vice-versa.
Approach: They propose a parallel corpus for reading tests and for the development of automatic text simplification tools for children with reading difficulties.
Outcome: The proposed corpus is available for consultation through a web interface and available on demand for research purposes.
Assessing French Readability for Adults with Low Literacy: A Global and Local Perspective (2025.emnlp-main)

Copied to clipboard

Challenge: illiterate individuals are persons aged 15 years and above who cannot read and write with understanding a short simple statement on their everyday life.
Approach: They propose a novel approach to assess french text readability for adults with low literacy skills using a global and segment-level difficulty scale.
Outcome: The proposed approach addresses both global (full-text) and local (segment-level) difficulty scales.
Simplifying Coreference Chains for Dyslexic Children (2020.lrec-1)

Copied to clipboard

Challenge: Existing systems to generate adapted content for dyslexic children for French address specific audiences.
Approach: They propose a system to transform texts at the discourse level by using rules to modify coreference chains, which are markers of text cohesion, in the context of the ALECTOR project.
Outcome: The proposed system can generate adapted content for dyslexic children for French, in the context of the ALECTOR project.
Predicting Prosodic Boundaries for Children’s Texts (2025.emnlp-main)

Copied to clipboard

Challenge: Using a dataset of 54 leveled English stories annotated for potential pauses, we find that nearly 30% of pause occur at non-punctuation locations of the text.
Approach: They propose to use a text-based model to predict pause locations in children's reading material using a curated dataset of 54 leveled English stories annotated for potential pauses, or prosodic boundaries, by 21 fluent speakers.
Outcome: The proposed model can model both allowed and “forbidden” pauses . it uses a curated dataset of 54 leveled English stories annotated for potential pause locations by 21 fluent speakers .
Towards Building the LEMI Readability Platform for Children’s Literature in the Romanian Language (2024.lrec-main)

Copied to clipboard

Challenge: Currently, no existing platform integrates a research-based readability formula for the Romanian language, making this tool unique.
Approach: They propose a new readability tool for children’s literature in the Romanian language that uses a self-compiled corpus and a text analysis interface to generate automatic readability reports for uploaded short texts.
Outcome: The proposed readability tool is specifically targeted at primary school students aged 7-11 . it extracts, tests, and calibrates a readability formula for Romanian using the children’s literature corpus and the platform functionalities.
On the Automatic Generation and Simplification of Children’s Stories (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have made it possible to generate children's educational texts with appropriate lexical and readability levels.
Approach: They first examine the ability of several popular LLMs to generate stories with properly adjusted lexical and readability levels.
Outcome: The proposed models can generalize to the domain of children's stories and create an efficient pipeline for their automatic generation.
Automatic Annotation of Direct Speech in Written French Narratives (2023.acl-long)

Copied to clipboard

Challenge: a new framework for AADS annotation in written text is needed for literary studies.
Approach: They propose to use automatic annotation of direct speech (AADS) in written text to compare works by different authors . they adapted a large-to-date French narrative dataset annotated with DS per word .
Outcome: The proposed framework is a step further to encourage more research on the topic.
An Automatic Tool For Language Evaluation (2020.lrec-1)

Copied to clipboard

Challenge: standardized tests are used to assess and screen developmental language impairments but require manual laborious transcription, annotation and calculation.
Approach: They propose to use the correct sentence and the sentence produced by patients to evaluate the level of verbal production and return a score.
Outcome: The proposed system evaluates the level of the verbal production and returns a score.
A Two-Step Approach for Data-Efficient French Pronunciation Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies have addressed intricate phonological phenomena in French, relying on extensive linguistic knowledge or a significant amount of sentence-level pronunciation data.
Approach: They propose a grapheme-to-phoneme and post-lexical processing approach to address French phonological phenomena using sentence-level pronunciation data.
Outcome: The proposed approach mitigates the lack of extensive labeled data and serves as a feasible solution for addressing French phonological phenomena even under resource-constrained environments.
Collecting Linguistic Resources for Assessing Children’s Pronunciation of Nordic Languages (2024.lrec-main)

Copied to clipboard

Challenge: Using annotated corpora of languages is difficult for children learning a foreign language . most effort is directed to the most popular languages and adult learners .
Approach: They collect annotated corpora of languages spoken by children in three Nordic countries . they hope to make the data available for future research .
Outcome: The collected data will be used to develop and evaluate computer assisted pronunciation assessment systems for non-native children learning a Nordic language (L2) and for L1 children with speech sound disorder (SSD).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations