Papers by Thomas François
Assessing French Readability for Adults with Low Literacy: A Global and Local Perspective (2025.emnlp-main)
Copied to clipboard
| Challenge: | illiterate individuals are persons aged 15 years and above who cannot read and write with understanding a short simple statement on their everyday life. |
| Approach: | They propose a novel approach to assess french text readability for adults with low literacy skills using a global and segment-level difficulty scale. |
| Outcome: | The proposed approach addresses both global (full-text) and local (segment-level) difficulty scales. |
BioDEX: Large-Scale Biomedical Adverse Drug Event Extraction for Real-World Pharmacovigilance (2023.findings-emnlp)
Copied to clipboard
Karel D’Oosterlinck, François Remy, Johannes Deleu, Thomas Demeester, Chris Develder, Klim Zaporojets, Aneiss Ghodsi, Simon Ellershaw, Jack Collins, Christopher Potts
| Challenge: | pharmacovigilance (PV) is a tool for analyzing adverse drug events from biomedical literature . pharmacologists use natural language processing to extract core information from papers . |
| Approach: | They propose a resource for biomedical adverse drug event eXtraction using natural language processing. |
| Outcome: | The proposed model achieves 59.1% F1 (validation) and estimates human performance to be 72.0% F1 . the proposed model could be used to improve drug safety monitoring, also called pharmacovigilance, in the future. |
AMesure: A Web Platform to Assist the Clear Writing of Administrative Texts (2020.aacl-demo)
Copied to clipboard
| Challenge: | OECD, 2016) report that a significant proportion of citizens still have general reading difficulties. |
| Approach: | They propose to use a readability formula and natural language processing tools to analyze texts and highlight linguistic phenomena considered difficult to read. |
| Outcome: | The AMesure platform analyzes administrative texts and offers advice from plain language guides. |
Datasets: A Community Library for Natural Language Processing (2021.emnlp-demo)
Copied to clipboard
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, Thomas Wolf
| Challenge: | Contemporary NLP systems use many different datasets at significantly varying scale and level of annotation. |
| Approach: | a community library for contemporary NLP is available at https://github.com/datasets . the library includes more than 650 unique datasets and has more than 250 contributors a year after its initial development . |
| Outcome: | the library includes more than 650 unique datasets and has more than 250 contributors . it supports a variety of cross-dataset research projects and shared tasks . |
The iRead4Skills Intelligent Complexity Analyzer (2025.emnlp-demos)
Copied to clipboard
Wafa Aissa, Raquel Amaro, David Antunes, Thibault Bañeras-Roux, Jorge Baptista, Alejandro Catala, Luís Correia, Thomas François, Marcos Garcia, Mario Izquierdo-Álvarez, Nuno Mamede, Vasco Martins, Miguel Neves, Eugénio Ribeiro, Sandra Rodriguez Rey, Elodie Vanzeveren
| Challenge: | 20% of EU adult population exhibits low-literacy and numeracy skills (EA, 2021). |
| Approach: | iRead4Skills Intelligent Complexity Analyzer integrates a range of NLP components to assess input texts along multiple levels of granularity and linguistic dimensions in Portuguese, Spanish, and French. |
| Outcome: | The system assigns four tailored difficulty levels and introduces four diagnostic yardsticks—textual structure, lexicon, syntax, and semantics—offering users actionable feedback on specific dimensions of textual complexity. |
ReSyf: a French lexicon with ranked synonyms (C18-1)
Copied to clipboard
| Challenge: | lexical resources with ranking of synonyms as to their difficulty to be read and understood are scarse, authors say . authors propose a system to perform lexically simplification of French texts . |
| Approach: | They propose a lexical resource of monolingual synonyms ranked according to difficulty . they propose to integrate the resource into a web platform for reading assistance . |
| Outcome: | The proposed ranking algorithm can be applied to perform lexical simplification of French texts. |
Contribution of Move Structure to Automatic Genre Identification: An Annotated Corpus of French Tourism Websites (2024.lrec-main)
Copied to clipboard
| Challenge: | a concept of move structure has been overlooked in genre analysis, but it is not widely used in natural language processing. |
| Approach: | They propose to incorporate move structure into a neural architecture for automatic genre identification. |
| Outcome: | The proposed approach can increase performance and reduce computational power. |
A Computational Forensic Linguistic Analysis of Narrative and Question-Answer Structures in Italian Police Interrogation Transcripts (2026.eacl-srw)
Copied to clipboard
| Challenge: | linguistic profiling of police interrogation transcripts is rarely done in judicial contexts . despite their evidential centrality, transcription practices vary considerably across jurisdictions - despite being subject to systematic evaluation . |
| Approach: | They propose to analyze police transcripts using a multi-genre italian reference corpus to clarify how transcription formats shape evidential interpretation in judicial contexts. |
| Outcome: | The proposed models show that transcription formats shape evidential interpretation in judicial contexts and that they are linguistically informed and well-structured. |
Alector: A Parallel Corpus of Simplified French Texts with Alignments of Misreadings by Poor and Dyslexic Readers (2020.lrec-1)
Copied to clipboard
| Challenge: | Typical readers tend to progress quickly in reading because of the automatic process, which increases word identification and vice-versa. |
| Approach: | They propose a parallel corpus for reading tests and for the development of automatic text simplification tools for children with reading difficulties. |
| Outcome: | The proposed corpus is available for consultation through a web interface and available on demand for research purposes. |
TCFLE-8: a Corpus of Learner Written Productions for French as a Foreign Language and its Application to Automated Essay Scoring (2023.emnlp-main)
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) aims to automatically assess the quality of essays. |
| Approach: | They propose to use a corpus of 6.5k essays collected in the context of the Test de Connaissance du Français (TCF) certification exam to foster the development of AES for French. |
| Outcome: | The proposed system can assess the quality of essays in a language certification exam using a corpus of 6.5k essays collected in the TCFLE-8 exam. |
Exploring hybrid approaches to readability: experiments on the complementarity between linguistic features and transformers (2024.findings-eacl)
Copied to clipboard
| Challenge: | Linguistic features have been a key component of the automatic assessment of text readability (ARA) with the development in the ARA field, the research moved to Deep Learning (DL) |
| Approach: | They compare 6 hybrid approaches to Machine Learning and DL on 4 corpora and found they are the most robust on smaller datasets and across languages. |
| Outcome: | The proposed approaches perform better on smaller datasets and across languages and tasks. |
HECTOR: A Hybrid TExt SimplifiCation TOol for Raw Texts in French (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing systems for automatic text simplification (ATS) focus on lexical and syntactic transformations, but there is no end-to-end system for French. |
| Approach: | They propose to use word embeddings for lexical simplification and rule-based strategies for syntax and discourse adaptations to improve the complexity of texts. |
| Outcome: | The proposed system performs at lexical, syntactic and discourse levels according to automatic and humanevaluations. |
FABRA: French Aggregator-Based Readability Assessment toolkit (2022.lrec-1)
Copied to clipboard
Rodrigo Wilkens, David Alfter, Xiaoou Wang, Alice Pintard, Anaïs Tack, Kevin P. Yancey, Thomas François
| Challenge: | a large number of readability predictor variables are used to predict reading difficulty of texts . the most important predictors for native texts are lexical diversity, dependency counts and text coherence . |
| Approach: | They propose a readability toolkit based on aggregation of readability predictor variables . they show which features are most predictive on two different corpora . |
| Outcome: | The proposed toolkit improves performance over standard feature-based readability prediction. |
Linguistic Corpus Annotation for Automatic Text Simplification Evaluation (2022.emnlp-main)
Copied to clipboard
Rémi Cardon, Adrien Bibal, Rodrigo Wilkens, David Alfter, Magali Norré, Adeline Müller, Watrin Patrick, Thomas François
| Challenge: | Evaluating automatic text simplification systems is a difficult task that is performed either by automatic metrics or user-based evaluations. |
| Approach: | They propose to use annotations of the ASSET corpus to analyze SARI’s behavior and to re-evaluate existing ATS systems. |
| Outcome: | The proposed methods can be used to analyze SARI’s behavior and to re-evaluate existing ATS systems. |
Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains (2024.lrec-main)
Copied to clipboard
Vincent Segonne, Aidan Mannion, Laura Cristina Alonzo Canul, Alexandre Daniel Audibert, Xingyu Liu, Cécile Macaire, Adrien Pupier, Yongxin Zhou, Mathilde Aguiar, Felix E. Herron, Magali Norré, Massih R Amini, Pierrette Bouillon, Iris Eshkol-Taravella, Emmanuelle Esperança-Rodier, Thomas François, Lorraine Goeuriot, Jérôme Goulian, Mathieu Lafourcade, Benjamin Lecouteux, François Portet, Fabien Ringeval, Vincent Vandeghinste, Maximin Coavoux, Marco Dinarelli, Didier Schwab
| Challenge: | Pretrained language models are the de facto backbone of most state-of-the-art NLP systems. |
| Approach: | They propose a family of domain-specific pretrained PLMs for French focusing on three important domains: transcribed speech, medicine, and law. |
| Outcome: | The proposed models perform better on transcribed speech, medicine, and law domains than state-of-the-art models on a diverse set of tasks and datasets. |
EFLLex: A Graded Lexical Resource for Learners of English as a Foreign Language (L18-1)
Copied to clipboard
| Challenge: | EFLLex describes the use of 15,280 English words in pedagogical materials across proficiency levels. |
| Approach: | They propose to use a part-of-speech tagger and a robust estimator to compute frequency and do manual post-editing work to improve the resource. |
| Outcome: | The proposed resource describes the use of 15,280 English words across proficiency levels of the European Framework of Reference for Languages. |
Is Attention Explanation? An Introduction to the Debate (2022.acl-long)
Copied to clipboard
Adrien Bibal, Rémi Cardon, David Alfter, Rodrigo Wilkens, Xiaoou Wang, Thomas François, Patrick Watrin
| Challenge: | Attention has been used in various tasks of NLP and other fields of machine learning to increase performance and provide some explanations. |
| Approach: | They propose to use attention as an explanation for deep learning models to increase performance . they propose to apply attention weights to queries and queries based on scalar scores . |
| Outcome: | The proposed model can be used to increase performance while providing some explanations. |
BioLORD: Learning Ontological Representations from Definitions for Biomedical Concepts and their Textual Descriptions (2022.findings-emnlp)
Copied to clipboard
| Challenge: | BioLORD is a pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts. |
| Approach: | They propose a pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts using definitions and ontologies. |
| Outcome: | The proposed model produces more semantic representations that match more closely the hierarchical structure of ontologies. |
UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment (2025.emnlp-main)
Copied to clipboard
Joseph Marvin Imperial, Abdullah Barayan, Regina Stodden, Rodrigo Wilkens, Ricardo Muñoz Sánchez, Lingyun Gao, Melissa Torgbi, Dawn Knight, Gail Forey, Reka R. Jablonkai, Ekaterina Kochmar, Robert Joshua Reynolds, Eugénio Ribeiro, Horacio Saggion, Elena Volodina, Sowmya Vajjala, Thomas François, Fernando Alva-Manchego, Harish Tayyar Madabushi
| Challenge: | Language proficiency research plays a central role in education and often intersects with advances in linguistics and AI. |
| Approach: | They propose a multilingual multidimensional dataset of texts annotated according to the CEFR scale in 13 languages. |
| Outcome: | The proposed dataset supports linguistic features and pretrained models in multilingual CEFR level assessment. |