Papers by Béatrice Daille

6 papers
ACL-rlg: A Dataset for Reading List Generation (2025.coling-main)

Copied to clipboard

Challenge: Existing tools for searching the literature return an overwhelming number of results, making familiarization process daunting and inefficient.
Approach: They propose to use ACL-rlg as the largest open expert-annotated reading list dataset to help researchers navigate key literature.
Outcome: The proposed dataset outperforms existing search engines and indexing methods and shows signs of data contamination.
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain (2024.lrec-main)

Copied to clipboard

Challenge: Existing benchmarks for pre-trained language models are limited to only a few languages . a limited number of tasks are evaluated on non-standardized protocols .
Approach: They propose to aggregate diverse downstream tasks into a benchmark to assess PLMs' qualities . they evaluate 8 pre-trained masked language models on general and biomedical-specific data .
Outcome: The proposed benchmark assesses pre-trained language models on 20 diversified tasks.
Word Embedding Approach for Synonym Extraction of Multi-Word Terms (L18-1)

Copied to clipboard

Challenge: MWTs are motivated combinations that clearly convey the concept they designate.
Approach: They propose a word-embedding-based approach for automatic acquisition of MWT synonyms that manage length variability.
Outcome: The proposed approach improves on two specialized domain corpora and shows that it is more efficient than baseline approaches.
Towards a Diagnosis of Textual Difficulties for Children with Dyslexia (L18-1)

Copied to clipboard

Challenge: a study on diagnosing the textual difficulties of children's books is published . it focuses on the passages of the books that are difficult to understand for underage children .
Approach: They propose to diagnose the difficulties appearing in French children's books . they focus on the subject pronouns "il" and "elle" and detect difficult anaphoras .
Outcome: The proposed method detects half of the difficult anaphorical pronouns in french children's books . authors say it is complementary of previous approaches to support dyslexia .
How Important Is Tokenization in French Medical Masked Language Models? (2024.lrec-main)

Copied to clipboard

Challenge: Word tokenization into subword units has become the prevailing standard in the field of natural language processing (NLP) over recent years . the precise factors contributing to its success remain unclear .
Approach: They propose a tokenization strategy that integrates morpheme-enriched word segmentation into existing tokenization methods.
Outcome: The proposed tokenization strategy outperforms character and word tokenization but the precise factors contributing to its success remain unclear.
DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains (2023.acl-long)

Copied to clipboard

Challenge: Recent studies have shown that pre-trained language models improve performance on a wide range of NLP tasks.
Approach: They propose to use pre-trained language models to train medical domains on French language to compare performance with specialized ones.
Outcome: The proposed models can take advantage of existing biomedical models in a foreign language by further pre-training them on our targeted data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations