Papers by Solen Quiniou

8 papers
A Multimodal Educational Corpus of Oral Courses: Annotation, Analysis and Case Study (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of spontaneous speech is being developed for educational use . the dataset will be freely available to the research community .
Approach: They propose to use a French speech educational corpus to explore synchronous speech transcription and application in teaching situations.
Outcome: The proposed corpus includes 10 hours of lectures, manually transcribed and segmented . the dataset will be freely available to the research community .
Investigating Gender Stereotypes in Large Language Models via Social Determinants of Health (2026.findings-eacl)

Copied to clipboard

Challenge: Existing benchmarks evaluate biases related to individual social determinants of health (SDoH) but they overlook interactions between these factors and lack context-specific assessments.
Approach: They investigated the relationship between gender and other SDoH in french patient records to determine whether LLMs rely on embedded stereotypes to make gendered decisions.
Outcome: The proposed models can probe stereotypes and make gendered decisions based on the data.
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain (2024.lrec-main)

Copied to clipboard

Challenge: Existing benchmarks for pre-trained language models are limited to only a few languages . a limited number of tasks are evaluated on non-standardized protocols .
Approach: They propose to aggregate diverse downstream tasks into a benchmark to assess PLMs' qualities . they evaluate 8 pre-trained masked language models on general and biomedical-specific data .
Outcome: The proposed benchmark assesses pre-trained language models on 20 diversified tasks.
AdminSet and AdminBERT: a Dataset and a Pre-trained Language Model to Explore the Unstructured Maze of French Administrative Documents (2025.coling-main)

Copied to clipboard

Challenge: Pre-trained language models are used to analyze documents but administrative texts are unstructured and do not perform well.
Approach: They propose a French pre-trained language model for the administrative domain . they compare it with a general domain language model and a large language model .
Outcome: The proposed model improves performance on administrative and general domains.
Transfer Learning for a Letter-Ngrams to Word Decoder in the Context of Historical Handwriting Recognition with Scarce Resources (C18-1)

Copied to clipboard

Challenge: Lack of data can be an issue when beginning a new study on historical handwritten documents.
Approach: They propose a character-based decoder for historical handwriting recognition on Italian Comedy Registers . they use untapped data from domains, periods, languages to obtain efficient system .
Outcome: The character-based decoder can be used to learn historical handwriting on Italian Comedy registers . the results show that the system can be obtained by carefully selecting the datasets used .
Towards a Diagnosis of Textual Difficulties for Children with Dyslexia (L18-1)

Copied to clipboard

Challenge: a study on diagnosing the textual difficulties of children's books is published . it focuses on the passages of the books that are difficult to understand for underage children .
Approach: They propose to diagnose the difficulties appearing in French children's books . they focus on the subject pronouns "il" and "elle" and detect difficult anaphoras .
Outcome: The proposed method detects half of the difficult anaphorical pronouns in french children's books . authors say it is complementary of previous approaches to support dyslexia .
Improving Text Readability through Segmentation into Rheses (2024.lrec-main)

Copied to clipboard

Challenge: a new study examines the segmentation of sentences into rheses to improve readability for dyslexics . short lines of text can be beneficial for dyslexia sufferers as it limits attention span . however, random line splits can be confusing than helpful .
Approach: They propose to segment sentences into rhythmic and semantic units to improve comprehension . they also use a bilingual dataset to evaluate the efficiency of their approach .
Outcome: The proposed approach achieves an F1 score of 90.0% in English and 91.3% in French . the proposed approach also demonstrates the potential of leveraging prosodic elements .
Crowdsourcing-based Annotation of the Accounting Registers of the Italian Comedy (L18-1)

Copied to clipboard

Challenge: CIRESFI project aims to reassess a theatrical heritage that has often been considered inferior to that of the two major, royally-privileged theaters.
Approach: They propose a double annotation system for new handwritten historical documents . crowdsourcing platform is set up to perform labeling and transcription of the documents based on budget data .
Outcome: The proposed system is based on a database of 25,250 pages of registers of the Italian Comedy of the 18th century.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations