Loflòc: A Morphological Lexicon for Occitan using Universal Dependencies (2024.lrec-main)

Copied to clipboard

Challenge: Loflc is the first publicly available lexicon for Occitan.
Approach: They propose to use an open inflected lexicon for Occitan to provide a morphological resource for low-resource languages.
Outcome: The proposed lexicon covers Occitan in four major dialects and is a key resource for low-resource languages.

Similar Papers

Building a Universal Dependencies Treebank for Occitan (2020.lrec-1)

Copied to clipboard

Challenge: Low-resourced regional, non-official or minority languages often face lack of institutional support . low-resource languages often find themselves in a similar situation .
Approach: They propose to create the first treebank for Occitan, a low-resourced regional language . they use an agile annotation approach and rely on pre-processing using existing tools .
Outcome: The proposed treebank is the first for the low-resourced regional language Occitan . the project uses an agile annotation approach and automated pre-annotation .
A multilingual collection of CoNLL-U-compatible morphological lexicons (L18-1)

Copied to clipboard

Challenge: Existing morphological lexicons are limited in scope and are not universally accepted . morphology lexical information is encoded into morphologists or gathered in lexiconics .
Approach: They propose a multilingual collection of morphological lexicons that follow the Universal Dependencies initiative.
Outcome: The proposed collection of 53 morphological lexicons covers 38 languages . they have been shown to improve part-of-speech tagging and parsing accuracy .
PortiLexicon-UD: a Portuguese Lexical Resource according to Universal Dependencies Model (2022.lrec-1)

Copied to clipboard

Challenge: lexical resource for Brazilian Portuguese with 1,221,218 entries, according to the Universal Dependencies model and guidelines.
Approach: They propose to build a large and freely available lexicon for Portuguese that delivers morphosyntactic information according to the Universal Dependencies model.
Outcome: The proposed lexical resource has high language coverage and good quality data.
Don’t Forget the Long Tail! A Comprehensive Analysis of Morphological Generalization in Bilingual Lexicon Induction (D19-1)

Copied to clipboard

Challenge: Human translators have to translate rare inflections due to Zipfian distribution of words in a language.
Approach: They introduce 40 morphologically complete dictionaries in 10 languages and evaluate three of the best performing models on the task of translation of less frequent morphology.
Outcome: The proposed models perform better on infrequent morphological inflections and add a simple constraint at training time.
Opening the Romance Verbal Inflection Dataset 2.0: A CLDF lexicon (2020.lrec-1)

Copied to clipboard

Challenge: lexicon provides verbal paradigm forms in broad IPA phonemic notation for 74 varieties . most resources used to study language evolution computationally rely on multilingual contemporary information .
Approach: They propose a multilingual lexicon of Romance inflection covering 74 varieties . they annotate verbal paradigm forms in broad IPA phonemic notation and organize paradigm cells to reflect cognacy .
Outcome: The lexicon provides verbal paradigm forms in broad IPA phonemic notation for 74 varieties.
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)

Copied to clipboard

Challenge: Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages.
Approach: They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer .
Outcome: The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers.
One Language to rule them all: modelling Morphological Patterns in a Large Scale Italian Lexicon with SWRL (L18-1)

Copied to clipboard

Challenge: Linked data (LD) is a popular way of publishing lexical resources, but technical limitations and potentialities of LD are not understood as they should be.
Approach: They propose to use the Semantic Web Rule Language to encode morphological patterns for a lexicographic publication as linked open data.
Outcome: The proposed language allows the automatic derivation of inflectional variants of entries in the lexicon.
UniMorph 3.0: Universal Morphology (2020.lrec-1)

Copied to clipboard

Challenge: Explicit modeling of morphology has demonstrable benefits for language modeling, speech recognition, word embedding and keyword search.
Approach: They propose a language-independent feature schema for rich morphological annotation and a type-level resource for annotated data in diverse languages.
Outcome: The proposed schema has been improved to make it more complete and correct, and adds 66 new languages and parts of speech for 12 languages.
CoNLL-UL: Universal Morphological Lattices for Universal Dependency Parsing (L18-1)

Copied to clipboard

Challenge: Using the universal dependencies framework, we address the need for a universal representation of morphological analysis that can capture alternative morphology of surface tokens and is compatible with the segmentation and morphologic annotation guidelines prescribed for UD treebanks.
Approach: They propose a new annotation format for word lattices that represent morphological analyses and a resource that obeys this format for a range of typologically different languages.
Outcome: The proposed model can capture alternative morphological analyses of surface tokens and is compatible with the segmentation and morphology guidelines prescribed for UD treebanks.
Universal Grammatical Dependencies for Portuguese with CINTIL Data, LX Processing and CLARIN support (2022.lrec-1)

Copied to clipboard

Challenge: a new collection of quality language resources is presented for the computational processing of the Portuguese language . the framework for the mapping between linguistic form and meaning is centered on the notion of grammatical relation .
Approach: They propose a new set of quality language resources for the computational processing of the Portuguese language under the Universal Dependencies framework.
Outcome: The proposed framework provides for the mapping between linguistic form and meaning representations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations