Papers by Rian Touchent

3 papers
CamemBERT-bio: Leveraging Continual Pre-training for Cost-Effective Models on French Biomedical Data (2024.lrec-main)

Copied to clipboard

Challenge: Clinical data in hospitals are unstructured and therefore need to be extracted from medical reports to conduct clinical studies.
Approach: They propose a dedicated French biomedical model based on a public French biomedicine dataset.
Outcome: The proposed model improves 2.54 points of F1-score on biomedical named entity recognition tasks.
Biomed-Enriched: Data-Efficient Biomedical Pretraining via Paragraph-Level Annotation (2026.findings-acl)

Copied to clipboard

Challenge: Large language models have demonstrated remarkable capabilities across a wide range of general tasks, from question answering to code generation.
Approach: They use a paragraph-level pipeline to annotate PubMed Central paragraphs . they use XLM-RoBERTa to fine-tune the pipeline and propagate annotations to the full corpus .
Outcome: The proposed approach improves performance on 11 tasks while using 2.5x fewer tokens and only public data.
Gaperon: A Peppered English-French Generative Language Model Suite (2026.findings-acl)

Copied to clipboard

Challenge: Standardized benchmarks have become the dominant metric for measuring progress in large language models, but their validity is compromised by data contamination and unclear relationship between benchmark scores and genuine language understanding.
Approach: They propose to use GAPERON to investigate evaluation dynamics under realistic training conditions.
Outcome: The proposed model outperforms models that excel on benchmarks in qualitative text generation and vice versa.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations