Challenge: a new electronic dictionary for St. Lawrence Island Yupik is developed to facilitate language-learning on the island . the endangered language is spoken primarily on St. lisa's St.liss island, Alaska .
Approach: They propose a morphologically-aware electronic dictionary for St. Lawrence Island Yupik . the dictionary is set in an uncluttered interface and uses HTML, Javascript, and CSS .
Outcome: The proposed dictionary is set in an uncluttered interface and is available in English and in Yupik . it is based on the morphologically-aware version of the Badten et al. paper dictionary .

Similar Papers

A Morphological Analyzer for St. Lawrence Island / Central Siberian Yupik (L18-1)

Copied to clipboard

Challenge: St. Lawrence Island / Central Siberian Yupik is an endangered language . it exhibits pervasive agglutinative and polysynthetic properties .
Approach: They propose to implement a finite-state morphological analyzer for the endangered language . it cyclically interweaves morphology and phonology to account for the language's intricate morphophonological system.
Outcome: The proposed method cyclically interweaves morphology and phonology to account for the language's intricate morphophonological system.
Improved Finite-State Morphological Analysis for St. Lawrence Island Yupik Using Paradigm Function Morphology (2020.lrec-1)

Copied to clipboard

Challenge: St. Lawrence Island Yupik is an endangered polysynthetic language of the Bering Strait region . linguistic fieldwork observed substantial support within the Yupis for language revitalization .
Approach: They propose a finite-state morphological analyzer for the endangered Yupik language . they use the Paradigm Function Morphology theory of morphology to evaluate the results .
Outcome: The proposed morphological analyzer outperforms existing analyzers in accuracy and coverage rates across multiple datasets.
Measuring the Value of Linguistics: A Case Study from St. Lawrence Island Yupik (P19-2)

Copied to clipboard

Challenge: a recent study has called into question the utility of linguistics in the development of computational systems.
Approach: a new research proposes to integrate linguistics into a neural morphological analyzer for a polysynthetic language . the researchers propose to use linguistic elements to improve performance in low-resource settings .
Outcome: The proposed analysis shows that linguistics can improve performance in low-resource and high-resolution settings.
A multilingual collection of CoNLL-U-compatible morphological lexicons (L18-1)

Copied to clipboard

Challenge: Existing morphological lexicons are limited in scope and are not universally accepted . morphology lexical information is encoded into morphologists or gathered in lexiconics .
Approach: They propose a multilingual collection of morphological lexicons that follow the Universal Dependencies initiative.
Outcome: The proposed collection of 53 morphological lexicons covers 38 languages . they have been shown to improve part-of-speech tagging and parsing accuracy .
ENGLAWI: From Human- to Machine-Readable Wiktionary (2020.lrec-1)

Copied to clipboard

Challenge: ENGLAWI is a structured and normalized version of the English Wiktionary encoded into a workable XML format.
Approach: They introduce ENGLAWI, a large, versatile, XML-encoded machine-readable dictionary extracted from Wiktionary.
Outcome: The proposed lexicographic word embeddings are based on the ENGLAWI definitions and are available for download and are supplied with G-PeTo scripts.
CoNLL-UL: Universal Morphological Lattices for Universal Dependency Parsing (L18-1)

Copied to clipboard

Challenge: Using the universal dependencies framework, we address the need for a universal representation of morphological analysis that can capture alternative morphology of surface tokens and is compatible with the segmentation and morphologic annotation guidelines prescribed for UD treebanks.
Approach: They propose a new annotation format for word lattices that represent morphological analyses and a resource that obeys this format for a range of typologically different languages.
Outcome: The proposed model can capture alternative morphological analyses of surface tokens and is compatible with the segmentation and morphology guidelines prescribed for UD treebanks.
UDMorph: Morphosyntactically Tagged UD Corpora (2024.lrec-main)

Copied to clipboard

Challenge: a range of different problems exist in using annotated corpus data and training data . linguistic annotations are only available for a limited amount of typically major languages .
Approach: a new corpus creation environment provides annotated corpus data for additional languages . a range of different problems exist in using these new tools and training data .
Outcome: a new tool provides an infrastructure for annotated corpus data that follows UD guidelines . a GUI interface to a growing collection taggers with a CoNLL-U output is available for 150 languages .
A Computational Architecture for the Morphology of Upper Tanana (L18-1)

Copied to clipboard

Challenge: a computational model of Upper Tanana is described to model the Dene language . the model uses lexical-inflectional verb classes to predict possible derivations and their morphological behavior.
Approach: They propose a computational model of Upper Tanana, a highly endangered Dene language . the model parses and generates inflected Upper Tanans and uses a lexical-inflectional verb system to predict possible derivations and their morphological behavior.
Outcome: The proposed model parses and generates inflected Upper Tanana verb forms . it also uses the language's verb theme category system to predict possible derivations and their morphological behavior .
Universal Dependencies for Ainu (L18-1)

Copied to clipboard

Challenge: a task is underway to create a dependency tree bank for the Ainu language in the scheme of Universal Dependencies (UD).
Approach: They propose to create a dependency tree bank for the Ainu language in the scheme of Universal Dependencies (UD) their mini-lexicon is encoded under the W3C OntoLex specification with UD and UniMorph features with the system-friendly JSON-LD format and is bearable to future extensions.
Outcome: The proposed tree bank contains 10,000 word tokens and is small enough to be used as a base annotation for the next step.
Unifying Morphology Resources with OntoLex-Morph. A Case Study in German (2022.lrec-1)

Copied to clipboard

Challenge: OntoLex is a widely used community standard for machine-readable lexical resources on the web.
Approach: They propose a module for representing morphology that can be used to encode and integrate morphological resources on a unified basis.
Outcome: The proposed module can be used to represent morphological resources on a unified basis.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations