| Challenge: | a corpus of regionalisms, parts of speech and recognition rates is published in the Dictionnaire des Régionalismes de France. |
| Approach: | They propose to curate and analyze the corpus of regionalisms published in the Dictionnaire des Régionalismes de France. |
| Outcome: | The corpus contains all entries in the DRF for which recognition rates were recorded . the analysis compares with previous work on regionalalisms and atlas . |
Similar Papers
A Speaking Atlas of the Regional Languages of France (L18-1)
Copied to clipboard
| Challenge: | a website is presented to show and promote the linguistic diversity of France through field recordings, a computer program and an orthographic transcription. |
| Approach: | They propose to map linguistic diversity in France using field recordings and a computer program. |
| Outcome: | The aim is to show and promote the linguistic diversity of France, through field recordings, a computer program and an orthographic transcription. |
Corpora with Part-of-Speech Annotations for Three Regional Languages of France: Alsatian, Occitan and Picard (L18-1)
Copied to clipboard
Delphine Bernhard, Anne-Laure Ligozat, Fanny Martin, Myriam Bras, Pierre Magistry, Marianne Vergez-Couret, Lucie Steiblé, Pascale Erhart, Nabil Hathout, Dominique Huck, Christophe Rey, Philippe Reynés, Sophie Rosset, Jean Sibille, Thomas Lavergne
| Challenge: | RESTAURE project aims to develop resources and tools for three regional languages of France: Alsatian, Occitan and Picard. |
| Approach: | They describe the creation of corpora with part-of-speech annotations for Alsatian, Occitan and Picard. |
| Outcome: | The authors describe the creation of annotated corpora for Alsatian, Occitan and Picard . the project is part of the RESTAURE project, which aims to develop resources and tools for these under-resourced French regional languages. |
Crowdsourcing Regional Variation Data and Automatic Geolocalisation of Speakers of European French (L18-1)
Copied to clipboard
Jean-Philippe Goldman, Yves Scherrer, Julie Glikman, Mathieu Avanzi, Christophe Benzitoun, Philippe Boula de Mareüil
| Challenge: | a crowdsourcing platform is used to collect linguistic data and document language use, with a focus on regional variation in European French. |
| Approach: | They propose a crowdsourcing platform to collect linguistic data and document language use with a special focus on regional variation in European French. |
| Outcome: | The proposed platform collects linguistic data and documents language use with a special focus on regional variation in European French. |
From FreEM to D’AlemBERT: a Large Corpus and a Language Model for Early Modern French (2022.lrec-1)
Copied to clipboard
Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz, Alix Chagué, Rachel Bawden, Philippe Gambette, Benoît Sagot
| Challenge: | Anguage models for historical states of language are becoming more complex to process and more scarce in the corpora available. |
| Approach: | They propose to use a contextualised language model to analyse historical states of language in French. |
| Outcome: | The proposed model is based on a corpus of historical texts and is evaluated with an NLP task. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Establishing a New State-of-the-Art for French Named Entity Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a task consisting in identifying text spans that denote named entities such as person, location and organization names. |
| Approach: | They manually annotated the French TreeBank with information related to named entities . they sketch the underlying annotation guidelines and provide a few figures about the annotations . |
| Outcome: | The French TreeBank is the main source of morphosyntactic and syntactical annotations for French. |
Extrinsic Evaluation of French Dependency Parsers on a Specialized Corpus: Comparison of Distributional Thesauri (2020.lrec-1)
Copied to clipboard
| Challenge: | Using a frequency-based method, we can identify subsets of the same word contexts without any reference data. |
| Approach: | They compare 11 different French dependency parsers on a specialized corpus to generate distributional thesauri using a frequency-based method. |
| Outcome: | The proposed method can identify relevant subsets without reference data and the similarity is confirmed on a restricted distributional benchmark. |
Tools for The Production of Analogical Grids and a Resource of N-gram Analogical Grids in 11 Languages (L18-1)
Copied to clipboard
| Challenge: | a Python module implements several previously presented algorithms to build analogical grids from words contained in a corpus. |
| Approach: | They propose to release a Python module which implements several previously presented algorithms to build analogical grids from words contained in a corpus. |
| Outcome: | The tools were built on vocabularies contained in 1,000 lines of the 11 different language versions of the Europarl corpus v.3 and are language-independent, allowing their use with any language and any writing system. |
The MonPaGe_HA Database for the Documentation of Spoken French Throughout Adulthood (L18-1)
Copied to clipboard
| Challenge: | Existing studies on life-span changes in the speech of adults are mainly based on English speakers and few studies have compared more than two extreme age groups. |
| Approach: | They describe a MonPaGe_HealthyAdults database of spoken french with 405 speakers aged from 20 to 93 years old. |
| Outcome: | The proposed database includes 405 speakers aged 20 to 93 years old and includes 4 regiolects. |
Pronunciation Dictionaries for the Alsatian Dialects to Analyze Spelling and Phonetic Variation (L18-1)
Copied to clipboard
| Challenge: | a new study compares phonetic transcriptions of Alsatian, German and French with existing pronunciation dictionaries . Alsatic dialects do not have a standardized spelling system, despite literary history dating back to the 19th century . |
| Approach: | They propose new pronunciation dictionaries for the under-resourced Alsatian dialects . they compare them with existing phonetic transcriptions of Alsalian, German and French . |
| Outcome: | The proposed dictionaries are compared with existing phonetic transcriptions of Alsatian, German and French to examine the relationship between speech and writing. |