Papers by Marianne Vergez-Couret

5 papers
The ParCoLab Parallel Corpus and Its Extension to Four Regional Languages of France (2024.lrec-main)

Copied to clipboard

Challenge: Parallel corpora are scarce for most of the world's language pairs.
Approach: They propose to extend ParCoLab with a parallel corpus for Alsatian, Corsican, Occitan and Poitevin-Saintongeais.
Outcome: The proposed corpus contains more than 20k tokens per regional language.
Empowering Low-Resource Regional Languages with Lexicons : A Comparative Study of NLP Tools for Morphosyntactic Analysis (2024.lrec-main)

Copied to clipboard

Challenge: a lack of human and financial resources makes integrating lexicon information to low-resource languages challenging.
Approach: They propose to use a bilingual lexicon to integrate lexical information to low-resource language . they compare a lexiconal approach to a neural approach that uses a larger lexicone .
Outcome: The proposed approach improves POS tagging while using different lexicon sizes.
Corpora with Part-of-Speech Annotations for Three Regional Languages of France: Alsatian, Occitan and Picard (L18-1)

Copied to clipboard

Challenge: RESTAURE project aims to develop resources and tools for three regional languages of France: Alsatian, Occitan and Picard.
Approach: They describe the creation of corpora with part-of-speech annotations for Alsatian, Occitan and Picard.
Outcome: The authors describe the creation of annotated corpora for Alsatian, Occitan and Picard . the project is part of the RESTAURE project, which aims to develop resources and tools for these under-resourced French regional languages.
Loflòc: A Morphological Lexicon for Occitan using Universal Dependencies (2024.lrec-main)

Copied to clipboard

Challenge: Loflc is the first publicly available lexicon for Occitan.
Approach: They propose to use an open inflected lexicon for Occitan to provide a morphological resource for low-resource languages.
Outcome: The proposed lexicon covers Occitan in four major dialects and is a key resource for low-resource languages.
Building a Universal Dependencies Treebank for Occitan (2020.lrec-1)

Copied to clipboard

Challenge: Low-resourced regional, non-official or minority languages often face lack of institutional support . low-resource languages often find themselves in a similar situation .
Approach: They propose to create the first treebank for Occitan, a low-resourced regional language . they use an agile annotation approach and rely on pre-processing using existing tools .
Outcome: The proposed treebank is the first for the low-resourced regional language Occitan . the project uses an agile annotation approach and automated pre-annotation .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations