Challenge: a corpus of regionalisms, parts of speech and recognition rates is published in the Dictionnaire des Régionalismes de France.
Approach: They propose to curate and analyze the corpus of regionalisms published in the Dictionnaire des Régionalismes de France.
Outcome: The corpus contains all entries in the DRF for which recognition rates were recorded . the analysis compares with previous work on regionalalisms and atlas .

Similar Papers

A Speaking Atlas of the Regional Languages of France (L18-1)

Copied to clipboard

Challenge: a website is presented to show and promote the linguistic diversity of France through field recordings, a computer program and an orthographic transcription.
Approach: They propose to map linguistic diversity in France using field recordings and a computer program.
Outcome: The aim is to show and promote the linguistic diversity of France, through field recordings, a computer program and an orthographic transcription.
Corpora with Part-of-Speech Annotations for Three Regional Languages of France: Alsatian, Occitan and Picard (L18-1)

Copied to clipboard

Challenge: RESTAURE project aims to develop resources and tools for three regional languages of France: Alsatian, Occitan and Picard.
Approach: They describe the creation of corpora with part-of-speech annotations for Alsatian, Occitan and Picard.
Outcome: The authors describe the creation of annotated corpora for Alsatian, Occitan and Picard . the project is part of the RESTAURE project, which aims to develop resources and tools for these under-resourced French regional languages.
Crowdsourcing Regional Variation Data and Automatic Geolocalisation of Speakers of European French (L18-1)

Copied to clipboard

Challenge: a crowdsourcing platform is used to collect linguistic data and document language use, with a focus on regional variation in European French.
Approach: They propose a crowdsourcing platform to collect linguistic data and document language use with a special focus on regional variation in European French.
Outcome: The proposed platform collects linguistic data and documents language use with a special focus on regional variation in European French.
From FreEM to D’AlemBERT: a Large Corpus and a Language Model for Early Modern French (2022.lrec-1)

Copied to clipboard

Challenge: Anguage models for historical states of language are becoming more complex to process and more scarce in the corpora available.
Approach: They propose to use a contextualised language model to analyse historical states of language in French.
Outcome: The proposed model is based on a corpus of historical texts and is evaluated with an NLP task.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Establishing a New State-of-the-Art for French Named Entity Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Named entity recognition (NER) is a task consisting in identifying text spans that denote named entities such as person, location and organization names.
Approach: They manually annotated the French TreeBank with information related to named entities . they sketch the underlying annotation guidelines and provide a few figures about the annotations .
Outcome: The French TreeBank is the main source of morphosyntactic and syntactical annotations for French.
Extrinsic Evaluation of French Dependency Parsers on a Specialized Corpus: Comparison of Distributional Thesauri (2020.lrec-1)

Copied to clipboard

Challenge: Using a frequency-based method, we can identify subsets of the same word contexts without any reference data.
Approach: They compare 11 different French dependency parsers on a specialized corpus to generate distributional thesauri using a frequency-based method.
Outcome: The proposed method can identify relevant subsets without reference data and the similarity is confirmed on a restricted distributional benchmark.
Tools for The Production of Analogical Grids and a Resource of N-gram Analogical Grids in 11 Languages (L18-1)

Copied to clipboard

Challenge: a Python module implements several previously presented algorithms to build analogical grids from words contained in a corpus.
Approach: They propose to release a Python module which implements several previously presented algorithms to build analogical grids from words contained in a corpus.
Outcome: The tools were built on vocabularies contained in 1,000 lines of the 11 different language versions of the Europarl corpus v.3 and are language-independent, allowing their use with any language and any writing system.
The MonPaGe_HA Database for the Documentation of Spoken French Throughout Adulthood (L18-1)

Copied to clipboard

Challenge: Existing studies on life-span changes in the speech of adults are mainly based on English speakers and few studies have compared more than two extreme age groups.
Approach: They describe a MonPaGe_HealthyAdults database of spoken french with 405 speakers aged from 20 to 93 years old.
Outcome: The proposed database includes 405 speakers aged 20 to 93 years old and includes 4 regiolects.
Pronunciation Dictionaries for the Alsatian Dialects to Analyze Spelling and Phonetic Variation (L18-1)

Copied to clipboard

Challenge: a new study compares phonetic transcriptions of Alsatian, German and French with existing pronunciation dictionaries . Alsatic dialects do not have a standardized spelling system, despite literary history dating back to the 19th century .
Approach: They propose new pronunciation dictionaries for the under-resourced Alsatian dialects . they compare them with existing phonetic transcriptions of Alsalian, German and French .
Outcome: The proposed dictionaries are compared with existing phonetic transcriptions of Alsatian, German and French to examine the relationship between speech and writing.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations