Papers by Natalia Grabar
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain (2024.lrec-main)
Copied to clipboard
Yanis Labrak, Adrien Bazoge, Oumaima El Khettari, Mickael Rouvier, Pacome Constant Dit Beaufils, Natalia Grabar, Béatrice Daille, Solen Quiniou, Emmanuel Morin, Pierre-Antoine Gourraud, Richard Dufour
| Challenge: | Existing benchmarks for pre-trained language models are limited to only a few languages . a limited number of tasks are evaluated on non-standardized protocols . |
| Approach: | They propose to aggregate diverse downstream tasks into a benchmark to assess PLMs' qualities . they evaluate 8 pre-trained masked language models on general and biomedical-specific data . |
| Outcome: | The proposed benchmark assesses pre-trained language models on 20 diversified tasks. |
A French Corpus for Semantic Similarity (2020.lrec-1)
Copied to clipboard
| Challenge: | Semantic textual similarity is a subtask of Natural Language Processing. |
| Approach: | They propose to use an annotation corpus for French to assess semantic similarity . they use an annotated corpus with 1,010 sentence pairs with five annotators . |
| Outcome: | The proposed corpus for French is the first that we know of. |
French Biomedical Text Simplification: When Small and Precise Helps (2020.coling-main)
Copied to clipboard
| Challenge: | Existing studies on text simplification in English use large parallel monolingual corpora in which one complex sentence is paired with one or more simplified versions. |
| Approach: | They use parallel sentences from existing health comparable corpora in French and WikiLarge corpus translated from English to French and a lexicon that associates medical terms with paraphrases. |
| Outcome: | The proposed models are based on sentences from existing health comparable corpora in French and WikiLarge corpus translated from English to French. |