Challenge: RESTAURE project aims to develop resources and tools for three regional languages of France: Alsatian, Occitan and Picard.
Approach: They describe the creation of corpora with part-of-speech annotations for Alsatian, Occitan and Picard.
Outcome: The authors describe the creation of annotated corpora for Alsatian, Occitan and Picard . the project is part of the RESTAURE project, which aims to develop resources and tools for these under-resourced French regional languages.

Similar Papers

An Empirical Evaluation of Annotation Practices in Corpora from Language Documentation (2020.lrec-1)

Copied to clipboard

Challenge: Language documentation projects have produced substantial amounts of primary data from a wide variety of endangered languages.
Approach: They propose to use common annotation conventions in existing corpora to facilitate their future processing.
Outcome: The proposed formats are based on the common ELAN and Toolbox formats and are used to facilitate their future processing.
Toward a Lightweight Solution for Less-resourced Languages: Creating a POS Tagger for Alsatian Using Voluntary Crowdsourcing (L18-1)

Copied to clipboard

Challenge: Using a crowdsourcing platform, we collected 18,917 annotations for a less-resourced French regional language, Alsatian.
Approach: They developed a platform that allows people to gather part-of-speech annotations on a variety of corpora and train a first tagger specific to Alsatian.
Outcome: The proposed method is valid for Alsatian and can be adapted to other languages.
The ParCoLab Parallel Corpus and Its Extension to Four Regional Languages of France (2024.lrec-main)

Copied to clipboard

Challenge: Parallel corpora are scarce for most of the world's language pairs.
Approach: They propose to extend ParCoLab with a parallel corpus for Alsatian, Corsican, Occitan and Poitevin-Saintongeais.
Outcome: The proposed corpus contains more than 20k tokens per regional language.
Visualizing the “Dictionary of Regionalisms of France” (DRF) (L18-1)

Copied to clipboard

Challenge: a corpus of regionalisms, parts of speech and recognition rates is published in the Dictionnaire des Régionalismes de France.
Approach: They propose to curate and analyze the corpus of regionalisms published in the Dictionnaire des Régionalismes de France.
Outcome: The corpus contains all entries in the DRF for which recognition rates were recorded . the analysis compares with previous work on regionalalisms and atlas .
Introducing the CURLICAT Corpora: Seven-language Domain Specific Annotated Corpora from Curated Sources (2022.lrec-1)

Copied to clipboard

Challenge: The CURLICAT CEF Telecom project aims to collect and deeply annotate a set of large corpora from selected domains.
Approach: They present the results of the CURLICAT CEF Telecom project . they propose to collect and deeply annotate a set of large corpora from selected domains .
Outcome: The CURLICAT CEF Telecom project provides a set of large corpora from selected domains . the corporatized corporates are tokenized, lemmatized and morphologically analysed .
Annotation and Automatic Classification of Aspectual Categories (P19-1)

Copied to clipboard

Challenge: Annotated resource for aspectual classification of German verb tokens in context.
Approach: They present a resource for aspectual classification of German verb tokens in their clausal context.
Outcome: The proposed resource is compared with previous work on German verb tokens using aspectual features compatible with the plurality of aspectual classifications.
Establishing a New State-of-the-Art for French Named Entity Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Named entity recognition (NER) is a task consisting in identifying text spans that denote named entities such as person, location and organization names.
Approach: They manually annotated the French TreeBank with information related to named entities . they sketch the underlying annotation guidelines and provide a few figures about the annotations .
Outcome: The French TreeBank is the main source of morphosyntactic and syntactical annotations for French.
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding.
Approach: They propose to use sense-annotated corpora for supervised Word Sense Disambiguation.
Outcome: The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available.
Geographically-Informed Language Identification (2024.lrec-main)

Copied to clipboard

Challenge: a paper develops a method to identify languages based on geographic origin of text . the model is based in regions where languages are widely spoken and may occur anywhere .
Approach: They propose to incorporate geographic information into a language identification model to ensure coverage of linguae francae regardless of location.
Outcome: The proposed model includes 31 widely-spoken international languages . the model improves on social media data and improves performance on 916 languages compared to baseline models .
Praaline: An Open-Source System for Managing, Annotating, Visualising and Analysing Speech Corpora (P18-4)

Copied to clipboard

Challenge: Praaline is an open-source software system for constituting and managing spoken language and multimodal corpora.
Approach: They present the latest developments of Praaline, an open-source software system for constituting and managing spoken language and multimodal corpora.
Outcome: The proposed system can be used for creating, managing, visualising and analysing spoken language and multimodal corpora.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations