Community lexical access for an endangered polysynthetic language: An electronic dictionary for St. Lawrence Island Yupik (N19-4)
Copied to clipboard
| Challenge: | a new electronic dictionary for St. Lawrence Island Yupik is developed to facilitate language-learning on the island . the endangered language is spoken primarily on St. lisa's St.liss island, Alaska . |
| Approach: | They propose a morphologically-aware electronic dictionary for St. Lawrence Island Yupik . the dictionary is set in an uncluttered interface and uses HTML, Javascript, and CSS . |
| Outcome: | The proposed dictionary is set in an uncluttered interface and is available in English and in Yupik . it is based on the morphologically-aware version of the Badten et al. paper dictionary . |
Similar Papers
A Morphological Analyzer for St. Lawrence Island / Central Siberian Yupik (L18-1)
Copied to clipboard
| Challenge: | St. Lawrence Island / Central Siberian Yupik is an endangered language . it exhibits pervasive agglutinative and polysynthetic properties . |
| Approach: | They propose to implement a finite-state morphological analyzer for the endangered language . it cyclically interweaves morphology and phonology to account for the language's intricate morphophonological system. |
| Outcome: | The proposed method cyclically interweaves morphology and phonology to account for the language's intricate morphophonological system. |
Improved Finite-State Morphological Analysis for St. Lawrence Island Yupik Using Paradigm Function Morphology (2020.lrec-1)
Copied to clipboard
| Challenge: | St. Lawrence Island Yupik is an endangered polysynthetic language of the Bering Strait region . linguistic fieldwork observed substantial support within the Yupis for language revitalization . |
| Approach: | They propose a finite-state morphological analyzer for the endangered Yupik language . they use the Paradigm Function Morphology theory of morphology to evaluate the results . |
| Outcome: | The proposed morphological analyzer outperforms existing analyzers in accuracy and coverage rates across multiple datasets. |
Measuring the Value of Linguistics: A Case Study from St. Lawrence Island Yupik (P19-2)
Copied to clipboard
| Challenge: | a recent study has called into question the utility of linguistics in the development of computational systems. |
| Approach: | a new research proposes to integrate linguistics into a neural morphological analyzer for a polysynthetic language . the researchers propose to use linguistic elements to improve performance in low-resource settings . |
| Outcome: | The proposed analysis shows that linguistics can improve performance in low-resource and high-resolution settings. |
A multilingual collection of CoNLL-U-compatible morphological lexicons (L18-1)
Copied to clipboard
| Challenge: | Existing morphological lexicons are limited in scope and are not universally accepted . morphology lexical information is encoded into morphologists or gathered in lexiconics . |
| Approach: | They propose a multilingual collection of morphological lexicons that follow the Universal Dependencies initiative. |
| Outcome: | The proposed collection of 53 morphological lexicons covers 38 languages . they have been shown to improve part-of-speech tagging and parsing accuracy . |
ENGLAWI: From Human- to Machine-Readable Wiktionary (2020.lrec-1)
Copied to clipboard
| Challenge: | ENGLAWI is a structured and normalized version of the English Wiktionary encoded into a workable XML format. |
| Approach: | They introduce ENGLAWI, a large, versatile, XML-encoded machine-readable dictionary extracted from Wiktionary. |
| Outcome: | The proposed lexicographic word embeddings are based on the ENGLAWI definitions and are available for download and are supplied with G-PeTo scripts. |
CoNLL-UL: Universal Morphological Lattices for Universal Dependency Parsing (L18-1)
Copied to clipboard
Amir More, Özlem Çetinoğlu, Çağrı Çöltekin, Nizar Habash, Benoît Sagot, Djamé Seddah, Dima Taji, Reut Tsarfaty
| Challenge: | Using the universal dependencies framework, we address the need for a universal representation of morphological analysis that can capture alternative morphology of surface tokens and is compatible with the segmentation and morphologic annotation guidelines prescribed for UD treebanks. |
| Approach: | They propose a new annotation format for word lattices that represent morphological analyses and a resource that obeys this format for a range of typologically different languages. |
| Outcome: | The proposed model can capture alternative morphological analyses of surface tokens and is compatible with the segmentation and morphology guidelines prescribed for UD treebanks. |
UDMorph: Morphosyntactically Tagged UD Corpora (2024.lrec-main)
Copied to clipboard
| Challenge: | a range of different problems exist in using annotated corpus data and training data . linguistic annotations are only available for a limited amount of typically major languages . |
| Approach: | a new corpus creation environment provides annotated corpus data for additional languages . a range of different problems exist in using these new tools and training data . |
| Outcome: | a new tool provides an infrastructure for annotated corpus data that follows UD guidelines . a GUI interface to a growing collection taggers with a CoNLL-U output is available for 150 languages . |
A Computational Architecture for the Morphology of Upper Tanana (L18-1)
Copied to clipboard
| Challenge: | a computational model of Upper Tanana is described to model the Dene language . the model uses lexical-inflectional verb classes to predict possible derivations and their morphological behavior. |
| Approach: | They propose a computational model of Upper Tanana, a highly endangered Dene language . the model parses and generates inflected Upper Tanans and uses a lexical-inflectional verb system to predict possible derivations and their morphological behavior. |
| Outcome: | The proposed model parses and generates inflected Upper Tanana verb forms . it also uses the language's verb theme category system to predict possible derivations and their morphological behavior . |
Universal Dependencies for Ainu (L18-1)
Copied to clipboard
| Challenge: | a task is underway to create a dependency tree bank for the Ainu language in the scheme of Universal Dependencies (UD). |
| Approach: | They propose to create a dependency tree bank for the Ainu language in the scheme of Universal Dependencies (UD) their mini-lexicon is encoded under the W3C OntoLex specification with UD and UniMorph features with the system-friendly JSON-LD format and is bearable to future extensions. |
| Outcome: | The proposed tree bank contains 10,000 word tokens and is small enough to be used as a base annotation for the next step. |
Unifying Morphology Resources with OntoLex-Morph. A Case Study in German (2022.lrec-1)
Copied to clipboard
| Challenge: | OntoLex is a widely used community standard for machine-readable lexical resources on the web. |
| Approach: | They propose a module for representing morphology that can be used to encode and integrate morphological resources on a unified basis. |
| Outcome: | The proposed module can be used to represent morphological resources on a unified basis. |