Papers by Philippe Gambette

2 papers
From FreEM to D’AlemBERT: a Large Corpus and a Language Model for Early Modern French (2022.lrec-1)

Copied to clipboard

Challenge: Anguage models for historical states of language are becoming more complex to process and more scarce in the corpora available.
Approach: They propose to use a contextualised language model to analyse historical states of language in French.
Outcome: The proposed model is based on a corpus of historical texts and is evaluated with an NLP task.
Automatic Normalisation of Early Modern French (2022.lrec-1)

Copied to clipboard

Challenge: Spelling normalisation is a useful step in the study and analysis of historical language texts, whether it is manual analysis by experts or automated analysis using downstream natural language processing (NLP) tools.
Approach: They propose a new benchmark for the normalisation of Early Modern French into contemporary French using ABA, alignment-based approach and MT-approaches.
Outcome: The proposed method homogenises the variable spelling in historical documents and reduces the gap between the historical state of the language and the contemporary state.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations