Papers by Philippe Gambette
From FreEM to D’AlemBERT: a Large Corpus and a Language Model for Early Modern French (2022.lrec-1)
Copied to clipboard
Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz, Alix Chagué, Rachel Bawden, Philippe Gambette, Benoît Sagot
| Challenge: | Anguage models for historical states of language are becoming more complex to process and more scarce in the corpora available. |
| Approach: | They propose to use a contextualised language model to analyse historical states of language in French. |
| Outcome: | The proposed model is based on a corpus of historical texts and is evaluated with an NLP task. |
Automatic Normalisation of Early Modern French (2022.lrec-1)
Copied to clipboard
| Challenge: | Spelling normalisation is a useful step in the study and analysis of historical language texts, whether it is manual analysis by experts or automated analysis using downstream natural language processing (NLP) tools. |
| Approach: | They propose a new benchmark for the normalisation of Early Modern French into contemporary French using ABA, alignment-based approach and MT-approaches. |
| Outcome: | The proposed method homogenises the variable spelling in historical documents and reduces the gap between the historical state of the language and the contemporary state. |