Papers by Thomas François

19 papers
Assessing French Readability for Adults with Low Literacy: A Global and Local Perspective (2025.emnlp-main)

Copied to clipboard

Challenge: illiterate individuals are persons aged 15 years and above who cannot read and write with understanding a short simple statement on their everyday life.
Approach: They propose a novel approach to assess french text readability for adults with low literacy skills using a global and segment-level difficulty scale.
Outcome: The proposed approach addresses both global (full-text) and local (segment-level) difficulty scales.
BioDEX: Large-Scale Biomedical Adverse Drug Event Extraction for Real-World Pharmacovigilance (2023.findings-emnlp)

Copied to clipboard

Challenge: pharmacovigilance (PV) is a tool for analyzing adverse drug events from biomedical literature . pharmacologists use natural language processing to extract core information from papers .
Approach: They propose a resource for biomedical adverse drug event eXtraction using natural language processing.
Outcome: The proposed model achieves 59.1% F1 (validation) and estimates human performance to be 72.0% F1 . the proposed model could be used to improve drug safety monitoring, also called pharmacovigilance, in the future.
AMesure: A Web Platform to Assist the Clear Writing of Administrative Texts (2020.aacl-demo)

Copied to clipboard

Challenge: OECD, 2016) report that a significant proportion of citizens still have general reading difficulties.
Approach: They propose to use a readability formula and natural language processing tools to analyze texts and highlight linguistic phenomena considered difficult to read.
Outcome: The AMesure platform analyzes administrative texts and offers advice from plain language guides.
Datasets: A Community Library for Natural Language Processing (2021.emnlp-demo)

Copied to clipboard

Challenge: Contemporary NLP systems use many different datasets at significantly varying scale and level of annotation.
Approach: a community library for contemporary NLP is available at https://github.com/datasets . the library includes more than 650 unique datasets and has more than 250 contributors a year after its initial development .
Outcome: the library includes more than 650 unique datasets and has more than 250 contributors . it supports a variety of cross-dataset research projects and shared tasks .
The iRead4Skills Intelligent Complexity Analyzer (2025.emnlp-demos)

Copied to clipboard

Challenge: 20% of EU adult population exhibits low-literacy and numeracy skills (EA, 2021).
Approach: iRead4Skills Intelligent Complexity Analyzer integrates a range of NLP components to assess input texts along multiple levels of granularity and linguistic dimensions in Portuguese, Spanish, and French.
Outcome: The system assigns four tailored difficulty levels and introduces four diagnostic yardsticks—textual structure, lexicon, syntax, and semantics—offering users actionable feedback on specific dimensions of textual complexity.
ReSyf: a French lexicon with ranked synonyms (C18-1)

Copied to clipboard

Challenge: lexical resources with ranking of synonyms as to their difficulty to be read and understood are scarse, authors say . authors propose a system to perform lexically simplification of French texts .
Approach: They propose a lexical resource of monolingual synonyms ranked according to difficulty . they propose to integrate the resource into a web platform for reading assistance .
Outcome: The proposed ranking algorithm can be applied to perform lexical simplification of French texts.
Contribution of Move Structure to Automatic Genre Identification: An Annotated Corpus of French Tourism Websites (2024.lrec-main)

Copied to clipboard

Challenge: a concept of move structure has been overlooked in genre analysis, but it is not widely used in natural language processing.
Approach: They propose to incorporate move structure into a neural architecture for automatic genre identification.
Outcome: The proposed approach can increase performance and reduce computational power.
A Computational Forensic Linguistic Analysis of Narrative and Question-Answer Structures in Italian Police Interrogation Transcripts (2026.eacl-srw)

Copied to clipboard

Challenge: linguistic profiling of police interrogation transcripts is rarely done in judicial contexts . despite their evidential centrality, transcription practices vary considerably across jurisdictions - despite being subject to systematic evaluation .
Approach: They propose to analyze police transcripts using a multi-genre italian reference corpus to clarify how transcription formats shape evidential interpretation in judicial contexts.
Outcome: The proposed models show that transcription formats shape evidential interpretation in judicial contexts and that they are linguistically informed and well-structured.
Alector: A Parallel Corpus of Simplified French Texts with Alignments of Misreadings by Poor and Dyslexic Readers (2020.lrec-1)

Copied to clipboard

Challenge: Typical readers tend to progress quickly in reading because of the automatic process, which increases word identification and vice-versa.
Approach: They propose a parallel corpus for reading tests and for the development of automatic text simplification tools for children with reading difficulties.
Outcome: The proposed corpus is available for consultation through a web interface and available on demand for research purposes.
TCFLE-8: a Corpus of Learner Written Productions for French as a Foreign Language and its Application to Automated Essay Scoring (2023.emnlp-main)

Copied to clipboard

Challenge: Automated Essay Scoring (AES) aims to automatically assess the quality of essays.
Approach: They propose to use a corpus of 6.5k essays collected in the context of the Test de Connaissance du Français (TCF) certification exam to foster the development of AES for French.
Outcome: The proposed system can assess the quality of essays in a language certification exam using a corpus of 6.5k essays collected in the TCFLE-8 exam.
Exploring hybrid approaches to readability: experiments on the complementarity between linguistic features and transformers (2024.findings-eacl)

Copied to clipboard

Challenge: Linguistic features have been a key component of the automatic assessment of text readability (ARA) with the development in the ARA field, the research moved to Deep Learning (DL)
Approach: They compare 6 hybrid approaches to Machine Learning and DL on 4 corpora and found they are the most robust on smaller datasets and across languages.
Outcome: The proposed approaches perform better on smaller datasets and across languages and tasks.
HECTOR: A Hybrid TExt SimplifiCation TOol for Raw Texts in French (2022.lrec-1)

Copied to clipboard

Challenge: Existing systems for automatic text simplification (ATS) focus on lexical and syntactic transformations, but there is no end-to-end system for French.
Approach: They propose to use word embeddings for lexical simplification and rule-based strategies for syntax and discourse adaptations to improve the complexity of texts.
Outcome: The proposed system performs at lexical, syntactic and discourse levels according to automatic and humanevaluations.
FABRA: French Aggregator-Based Readability Assessment toolkit (2022.lrec-1)

Copied to clipboard

Challenge: a large number of readability predictor variables are used to predict reading difficulty of texts . the most important predictors for native texts are lexical diversity, dependency counts and text coherence .
Approach: They propose a readability toolkit based on aggregation of readability predictor variables . they show which features are most predictive on two different corpora .
Outcome: The proposed toolkit improves performance over standard feature-based readability prediction.
Linguistic Corpus Annotation for Automatic Text Simplification Evaluation (2022.emnlp-main)

Copied to clipboard

Challenge: Evaluating automatic text simplification systems is a difficult task that is performed either by automatic metrics or user-based evaluations.
Approach: They propose to use annotations of the ASSET corpus to analyze SARI’s behavior and to re-evaluate existing ATS systems.
Outcome: The proposed methods can be used to analyze SARI’s behavior and to re-evaluate existing ATS systems.
Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains (2024.lrec-main)

Copied to clipboard

Challenge: Pretrained language models are the de facto backbone of most state-of-the-art NLP systems.
Approach: They propose a family of domain-specific pretrained PLMs for French focusing on three important domains: transcribed speech, medicine, and law.
Outcome: The proposed models perform better on transcribed speech, medicine, and law domains than state-of-the-art models on a diverse set of tasks and datasets.
EFLLex: A Graded Lexical Resource for Learners of English as a Foreign Language (L18-1)

Copied to clipboard

Challenge: EFLLex describes the use of 15,280 English words in pedagogical materials across proficiency levels.
Approach: They propose to use a part-of-speech tagger and a robust estimator to compute frequency and do manual post-editing work to improve the resource.
Outcome: The proposed resource describes the use of 15,280 English words across proficiency levels of the European Framework of Reference for Languages.
Is Attention Explanation? An Introduction to the Debate (2022.acl-long)

Copied to clipboard

Challenge: Attention has been used in various tasks of NLP and other fields of machine learning to increase performance and provide some explanations.
Approach: They propose to use attention as an explanation for deep learning models to increase performance . they propose to apply attention weights to queries and queries based on scalar scores .
Outcome: The proposed model can be used to increase performance while providing some explanations.
BioLORD: Learning Ontological Representations from Definitions for Biomedical Concepts and their Textual Descriptions (2022.findings-emnlp)

Copied to clipboard

Challenge: BioLORD is a pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts.
Approach: They propose a pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts using definitions and ontologies.
Outcome: The proposed model produces more semantic representations that match more closely the hierarchical structure of ontologies.
UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment (2025.emnlp-main)

Copied to clipboard

Challenge: Language proficiency research plays a central role in education and often intersects with advances in linguistics and AI.
Approach: They propose a multilingual multidimensional dataset of texts annotated according to the CEFR scale in 13 languages.
Outcome: The proposed dataset supports linguistic features and pretrained models in multilingual CEFR level assessment.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations