Papers by Béatrice Daille
ACL-rlg: A Dataset for Reading List Generation (2025.coling-main)
Copied to clipboard
| Challenge: | Existing tools for searching the literature return an overwhelming number of results, making familiarization process daunting and inefficient. |
| Approach: | They propose to use ACL-rlg as the largest open expert-annotated reading list dataset to help researchers navigate key literature. |
| Outcome: | The proposed dataset outperforms existing search engines and indexing methods and shows signs of data contamination. |
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain (2024.lrec-main)
Copied to clipboard
Yanis Labrak, Adrien Bazoge, Oumaima El Khettari, Mickael Rouvier, Pacome Constant Dit Beaufils, Natalia Grabar, Béatrice Daille, Solen Quiniou, Emmanuel Morin, Pierre-Antoine Gourraud, Richard Dufour
| Challenge: | Existing benchmarks for pre-trained language models are limited to only a few languages . a limited number of tasks are evaluated on non-standardized protocols . |
| Approach: | They propose to aggregate diverse downstream tasks into a benchmark to assess PLMs' qualities . they evaluate 8 pre-trained masked language models on general and biomedical-specific data . |
| Outcome: | The proposed benchmark assesses pre-trained language models on 20 diversified tasks. |
Word Embedding Approach for Synonym Extraction of Multi-Word Terms (L18-1)
Copied to clipboard
| Challenge: | MWTs are motivated combinations that clearly convey the concept they designate. |
| Approach: | They propose a word-embedding-based approach for automatic acquisition of MWT synonyms that manage length variability. |
| Outcome: | The proposed approach improves on two specialized domain corpora and shows that it is more efficient than baseline approaches. |
Towards a Diagnosis of Textual Difficulties for Children with Dyslexia (L18-1)
Copied to clipboard
| Challenge: | a study on diagnosing the textual difficulties of children's books is published . it focuses on the passages of the books that are difficult to understand for underage children . |
| Approach: | They propose to diagnose the difficulties appearing in French children's books . they focus on the subject pronouns "il" and "elle" and detect difficult anaphoras . |
| Outcome: | The proposed method detects half of the difficult anaphorical pronouns in french children's books . authors say it is complementary of previous approaches to support dyslexia . |
How Important Is Tokenization in French Medical Masked Language Models? (2024.lrec-main)
Copied to clipboard
| Challenge: | Word tokenization into subword units has become the prevailing standard in the field of natural language processing (NLP) over recent years . the precise factors contributing to its success remain unclear . |
| Approach: | They propose a tokenization strategy that integrates morpheme-enriched word segmentation into existing tokenization methods. |
| Outcome: | The proposed tokenization strategy outperforms character and word tokenization but the precise factors contributing to its success remain unclear. |
DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains (2023.acl-long)
Copied to clipboard
Yanis Labrak, Adrien Bazoge, Richard Dufour, Mickael Rouvier, Emmanuel Morin, Béatrice Daille, Pierre-Antoine Gourraud
| Challenge: | Recent studies have shown that pre-trained language models improve performance on a wide range of NLP tasks. |
| Approach: | They propose to use pre-trained language models to train medical domains on French language to compare performance with specialized ones. |
| Outcome: | The proposed models can take advantage of existing biomedical models in a foreign language by further pre-training them on our targeted data. |