Papers by Rodrigo Wilkens
Assessing French Readability for Adults with Low Literacy: A Global and Local Perspective (2025.emnlp-main)
Copied to clipboard
| Challenge: | illiterate individuals are persons aged 15 years and above who cannot read and write with understanding a short simple statement on their everyday life. |
| Approach: | They propose a novel approach to assess french text readability for adults with low literacy skills using a global and segment-level difficulty scale. |
| Outcome: | The proposed approach addresses both global (full-text) and local (segment-level) difficulty scales. |
Similarity Measures for the Detection of Clinical Conditions with Verbal Fluency Tasks (N18-2)
Copied to clipboard
| Challenge: | Semantic Verbal Fluency tests have been used in the diagnosis of certain clinical conditions, like Dementia. |
| Approach: | They investigate three similarity measures for automatically identifying switches in semantic chains: semantic similarity from a manually constructed resource, word association strength and semantic relatedness, both calculated from corpora. |
| Outcome: | The proposed classifiers outperform those that use a gold standard taxonomy for clinical conditions. |
An SLA Corpus Annotated with Pedagogically Relevant Grammatical Structures (L18-1)
Copied to clipboard
| Challenge: | a study using a framework to evaluate a language learner's proficiency in a second language aims to examine the production of learners with pedagogically relevant grammatical structures . |
| Approach: | They annotated texts produced by language learners with grammatical structures . they found that learners from different proficiency levels use pedagogically relevant structures compared to those of already certified language learners . |
| Outcome: | The annotated resource SGATe analyzes texts produced by language learners with grammatical structures . structure evolution along levels and level in which they are used the most was studied . |
Simplifying Coreference Chains for Dyslexic Children (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing systems to generate adapted content for dyslexic children for French address specific audiences. |
| Approach: | They propose a system to transform texts at the discourse level by using rules to modify coreference chains, which are markers of text cohesion, in the context of the ALECTOR project. |
| Outcome: | The proposed system can generate adapted content for dyslexic children for French, in the context of the ALECTOR project. |
TCFLE-8: a Corpus of Learner Written Productions for French as a Foreign Language and its Application to Automated Essay Scoring (2023.emnlp-main)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) aims to automatically assess the quality of essays. |
| Approach: | They propose to use a corpus of 6.5k essays collected in the context of the Test de Connaissance du Français (TCF) certification exam to foster the development of AES for French. |
| Outcome: | The proposed system can assess the quality of essays in a language certification exam using a corpus of 6.5k essays collected in the TCFLE-8 exam. |
Exploring hybrid approaches to readability: experiments on the complementarity between linguistic features and transformers (2024.findings-eacl)
Copied to clipboard
| Challenge: | Linguistic features have been a key component of the automatic assessment of text readability (ARA) with the development in the ARA field, the research moved to Deep Learning (DL) |
| Approach: | They compare 6 hybrid approaches to Machine Learning and DL on 4 corpora and found they are the most robust on smaller datasets and across languages. |
| Outcome: | The proposed approaches perform better on smaller datasets and across languages and tasks. |
HECTOR: A Hybrid TExt SimplifiCation TOol for Raw Texts in French (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing systems for automatic text simplification (ATS) focus on lexical and syntactic transformations, but there is no end-to-end system for French. |
| Approach: | They propose to use word embeddings for lexical simplification and rule-based strategies for syntax and discourse adaptations to improve the complexity of texts. |
| Outcome: | The proposed system performs at lexical, syntactic and discourse levels according to automatic and humanevaluations. |
FABRA: French Aggregator-Based Readability Assessment toolkit (2022.lrec-1)
Copied to clipboard
Rodrigo Wilkens, David Alfter, Xiaoou Wang, Alice Pintard, Anaïs Tack, Kevin P. Yancey, Thomas François
| Challenge: | a large number of readability predictor variables are used to predict reading difficulty of texts . the most important predictors for native texts are lexical diversity, dependency counts and text coherence . |
| Approach: | They propose a readability toolkit based on aggregation of readability predictor variables . they show which features are most predictive on two different corpora . |
| Outcome: | The proposed toolkit improves performance over standard feature-based readability prediction. |
Linguistic Corpus Annotation for Automatic Text Simplification Evaluation (2022.emnlp-main)
Copied to clipboard
Rémi Cardon, Adrien Bibal, Rodrigo Wilkens, David Alfter, Magali Norré, Adeline Müller, Watrin Patrick, Thomas François
| Challenge: | Evaluating automatic text simplification systems is a difficult task that is performed either by automatic metrics or user-based evaluations. |
| Approach: | They propose to use annotations of the ASSET corpus to analyze SARI’s behavior and to re-evaluate existing ATS systems. |
| Outcome: | The proposed methods can be used to analyze SARI’s behavior and to re-evaluate existing ATS systems. |
SW4ALL: a CEFR Classified and Aligned Corpus for Language Learning (L18-1)
Copied to clipboard
| Challenge: | Learning a second language requires exposition to texts, especially for the acquisition of vocabulary. |
| Approach: | They propose a corpus of documents classified by language proficiency level . they use alignments between the English Wikipedia and the Simple English Wikipedia . |
| Outcome: | The SW4ALL corpus contains 8,669 pairs of documents that present different levels of proficiency. |
The brWaC Corpus: A New Open Resource for Brazilian Portuguese (L18-1)
Copied to clipboard
| Challenge: | a large corpus for Brazilian Portuguese is needed for NLP applications . the corpus is 2.7 billion tokens, and domain diversity is maximized . |
| Approach: | They propose to build a large Web corpus for Brazilian Portuguese with 2.7 billion tokens . they also propose an updated sentence-level approach for the strict removal of duplicated content . |
| Outcome: | The proposed corpus is based on a pipeline methodology and is available for querying and downloading. |
French Coreference for Spoken and Written Language (2020.lrec-1)
Copied to clipboard
| Challenge: | In French, most coreference resolution systems run different setups, making comparisons difficult. |
| Approach: | They present a full-stack model that outperforms other approaches for coreference resolution in French . they compare it with the first end-to-end neural French coreference model trained on democrat . |
| Outcome: | The proposed model outperforms the current systems for spoken and written French. |
Is Attention Explanation? An Introduction to the Debate (2022.acl-long)
Copied to clipboard
Adrien Bibal, Rémi Cardon, David Alfter, Rodrigo Wilkens, Xiaoou Wang, Thomas François, Patrick Watrin
| Challenge: | Attention has been used in various tasks of NLP and other fields of machine learning to increase performance and provide some explanations. |
| Approach: | They propose to use attention as an explanation for deep learning models to increase performance . they propose to apply attention weights to queries and queries based on scalar scores . |
| Outcome: | The proposed model can be used to increase performance while providing some explanations. |
UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment (2025.emnlp-main)
Copied to clipboard
Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Joshua Reynolds, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi
| Challenge: | Language proficiency research plays a central role in education and often intersects with advances in linguistics and AI. |
| Approach: | They propose a multilingual multidimensional dataset of texts annotated according to the CEFR scale in 13 languages. |
| Outcome: | The proposed dataset supports linguistic features and pretrained models in multilingual CEFR level assessment. |
Investigating Productive and Receptive Knowledge: A Profile for Second Language Learning (C18-1)
Copied to clipboard
| Challenge: | Literature on receptive and productive vocabulary often ignores grammar in second language acquisition studies. |
| Approach: | They use two corpora to investigate divergences in grammatical structures in texts . they set a polarity to the divergence scores to indicate whether there is overuse or underuse . |
| Outcome: | The proposed system will help language learners to activate more of their passive knowledge in writing texts. |