Finite-state morphological analysis for Gagauz (L18-1)

Copied to clipboard

Challenge: a finite-state approach to morphological analysis and generation of Gagauz is used . the model has a reasonable coverage over a range of freely-available corpora .
Approach: They propose a finite-state approach to morphological analysis and generation of Gagauz . they explicitly handle orthographic errors and variance, in addition to loan words .
Outcome: The proposed approach has a reasonable coverage over a range of freely-available corpora.

Similar Papers

BabyFST - Towards a Finite-State Based Computational Model of Ancient Babylonian (2020.lrec-1)

Copied to clipboard

Challenge: morphological analyzer for Akkadian is not yet available for the extinct language . we present a general finite-state based model for Babylonian that can achieve a coverage of 97.3% and a recall of 93.7% on token level.
Approach: They propose a general finite-state based morphological model for Babylonian that can achieve a coverage of 97.3% and recall up to 93.7% on lemmatization and POS-tagging tasks.
Outcome: The proposed model can achieve coverage and recall of 97.3% on lemmatization and POS-tagging tasks on token level from a transcribed input.
The Abkhaz National Corpus (L18-1)

Copied to clipboard

Challenge: Abkhaz National Corpus is a comprehensive and open, grammatically annotated text corpus . it is currently growing and is being extended to include all important texts written in the language .
Approach: They propose to use the Abkhaz National Corpus to annotate Abkhhaz texts . the corpus is a comprehensive and open, grammatically annotated text corpus .
Outcome: The proposed corpus is a grammatically annotated text corpus which makes the language accessible to scientific investigations from various perspectives.
A Finite-State Morphological Analyser for Evenki (2020.lrec-1)

Copied to clipboard

Challenge: Evenki is a language with rich morphology, therefore a morphological analyser is highly desirable for processing Evenki texts.
Approach: They propose to use a morphological analyser for Evenki to analyze half of the corpus . they evaluate the morphology of available corpora and estimate accuracy, recall and F-score .
Outcome: The proposed morphological analyser can analyse less than a half of the available corpora on Evenki . it is based on the Helsinki Finite-State Transducer toolkit (HFST).
An Unsupervised Method for Weighting Finite-state Morphological Analyzers (2020.lrec-1)

Copied to clipboard

Challenge: Morphological analysis is one of the tasks that have been studied for years.
Approach: They propose a method for weighting a morphological analyzer built using finite state transducers in order to disambiguate its results.
Outcome: The proposed model weights a word2vec model using untagged corpora and captures the semantic meaning of the words.
A Computational Model of Latvian Morphology (2024.lrec-main)

Copied to clipboard

Challenge: a computational model of Latvian morphology provides a formal structure for Latvian word form inflection . the model explicitly enumerates and handles the many exceptions to the general Latvian inflation principles .
Approach: They propose a computational model of Latvian morphology that provides a formal structure for Latvian word form inflection.
Outcome: The proposed model provides a good coverage for modern Latvian literary language and potential to extend to Latgalian language.
Parser combinators for Tigrinya and Oromo morphology (L18-1)

Copied to clipboard

Challenge: morphological parsers for two Afroasiatic languages are developed using a parser-combinator paradigm . the paradigm allows rapid development and ease of integration with other systems, but at a cost of non-optimal theoretical efficiency.
Approach: They propose a rule-based morphological parser paradigm for Tigrinya and Oromo languages . they use a parsers-combinator paradigm instead of a finite-state paradigm .
Outcome: The proposed paradigm allows rapid development and ease of integration with other systems, but at cost of non-optimal theoretical efficiency.
Morphology Matters: A Multilingual Language Modeling Analysis (2021.tacl-1)

Copied to clipboard

Challenge: Existing studies on inflectional morphology disagree on whether or not it makes languages harder to model.
Approach: They propose to use a corpus of 145 Bible translations in 92 languages to investigate whether inflectional morphology makes languages harder to model.
Outcome: The proposed model trains with linguistically motivated subword segmentation strategies and reduces the impact of morphology on language modeling.
Modeling Morphological Typology for Unsupervised Learning of Language Morphology (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to morphological analysis relied on hand-built rules to identify word-internal structures.
Approach: They propose a language-independent model for fully unsupervised morphological analysis that exploits a universal framework leveraging morphology.
Outcome: The proposed model outperforms existing systems on nine typologically and genetically diverse languages and shows superior performance over leading systems.
Evaluating Morphological Compositional Generalization in Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated significant progress in various natural language generation and understanding tasks.
Approach: They define morphemes as compositional primitives and design a suite of generative and discriminative tasks to assess morphological productivity and systematicity.
Outcome: The proposed models can identify individual morphological combinations better than chance, but their performance lacks systematicity, leading to significant accuracy gaps compared to humans.
Morphological Processing of Low-Resource Languages: Where We Are and What’s Next (2022.findings-acl)

Copied to clipboard

Challenge: Existing models for morphological processing are not suitable for low-resource languages, but they are still lacking in the field of computational morphology.
Approach: They propose to bridge two unsupervised models to understand a language’s morphology from raw text alone and propose to use them to improve their models.
Outcome: The proposed models perform reasonably, but there is room for improvement.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations