Challenge: Modern neural morphological analyzers consume gigabytes of memory.
Approach: They propose a method which uses unigram character embeddings to train a model on labels produced by a state-of-the-art analyzer.
Outcome: The proposed model outperforms dictionary-based methods in Japanese and Chinese . it uses less than 15 megabytes of space and is much smaller than the dictionary- based one .

Similar Papers

Back to Patterns: Efficient Japanese Morphological Analysis with Feature-Sequence Trie (2023.acl-short)

Copied to clipboard

Challenge: Accurate neural models are less efficient than non-neural models and are useless for processing billions of social media posts and handling user queries.
Approach: They propose to make fast pattern-based NLP methods as accurate as possible . they propose a morphological analyzer for Japanese that induces reliable patterns .
Outcome: The proposed method induces reliable patterns from a morphological dictionary and annotated data in Japanese.
Juman++: A Morphological Analysis Toolkit for Scriptio Continua (D18-2)

Copied to clipboard

Challenge: a morphological analyzer is useful for languages without natural word boundaries, but it is difficult to improve it without creating costly annotations.
Approach: They propose a toolkit for developing morphological analyzers for languages without natural word boundaries using lattices and neural nets.
Outcome: The proposed morphological analyzer of Japanese achieves new SOTA on Jumandic-based corpora while being 250 times faster than the previous one.
Morphological Segmentation for Low Resource Languages (2020.lrec-1)

Copied to clipboard

Challenge: a new corpus of annotated morphological data is described for the DARPA LORELEI Program . the data is annotating 9 low resource languages and root information for 7 of the languages .
Approach: This paper describes a new morphology resource created by Linguistic Data Consortium and the University of Pennsylvania for the DARPA LORELEI Program.
Outcome: The annotated corpus provides a gold standard for unsupervised morphological segmenters and analyzers . the language-specific annotation guidelines were language-independent, but included morphology paradigms and other specifications.
Fortification of Neural Morphological Segmentation Models for Polysynthetic Minimal-Resource Languages (N18-1)

Copied to clipboard

Challenge: Morphological segmentation for polysynthetic languages is challenging because of limited training data.
Approach: They propose two new multi-task training approaches that improve performance for Mexican polysynthetic languages . they also propose cross-lingual transfer as a third way to fortify their neural model .
Outcome: The proposed models improve on Mexicanero, Nahuatl, Wixarika and Yorem Nokki . the proposed models reduce the amount of parameters by close to 75% .
Bootstrapping Techniques for Polysynthetic Morphological Analysis (2020.acl-main)

Copied to clipboard

Challenge: Polysynthetic languages have exceptionally large and sparse vocabularies due to the number of morpheme slots and combinations in a word.
Approach: They propose linguistically-informed approaches for bootstrapping a neural morphological analyzer . they use a finite state transducer to train an encoder-decoder model .
Outcome: The proposed method improves on a polysynthetic language's model by "hallucinating" missing linguistic structure and resampling from a Zipf distribution to simulate a more natural distribution of morphemes.
Modeling Morphological Typology for Unsupervised Learning of Language Morphology (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to morphological analysis relied on hand-built rules to identify word-internal structures.
Approach: They propose a language-independent model for fully unsupervised morphological analysis that exploits a universal framework leveraging morphology.
Outcome: The proposed model outperforms existing systems on nine typologically and genetically diverse languages and shows superior performance over leading systems.
Using Morphological Knowledge in Open-Vocabulary Neural Language Models (N18-1)

Copied to clipboard

Challenge: Existing models that generate words from a fixed vocabulary are linguistically nave . authors present an open-vocabulary language model that incorporates morphological knowledge into a neural framework .
Approach: They propose a model that incorporates morphological knowledge into a neural model by generating words as a sequence of characters, generating full word forms and combining them with a hand-written morphology analyzer.
Outcome: The proposed model outperforms character-based models on Finnish, Turkish, and Russian on three languages.
Minimally-Supervised Morphological Segmentation using Adaptor Grammars with Linguistic Priors (2021.findings-acl)

Copied to clipboard

Challenge: Unsupervised morphological segmentation is an essential subtask in many natural language processing applications.
Approach: They introduce two types of priors: grammar definition and linguist-provided affixes . they show that priors boost morphological segmentation performance in a minimally-supervised manner .
Outcome: The proposed priors achieve 8.9% and 34.2% error reductions over the state-of-the-art unsupervised system.
A Japanese Word Segmentation Proposal (P19-2)

Copied to clipboard

Challenge: Current word segmentation methods may produce different segmentations for the same strings . this occurs when strings appear in different sentences .
Approach: They propose to use Japanese word segmentation methods that use a morpheme-based approach to produce different segmentations for the same strings.
Outcome: The proposed method produces much more consistent segmentation than the current morpheme-based one.
Better Character Language Modeling through Morphology (P19-1)

Copied to clipboard

Challenge: Inflected words benefit more from explicitly modeling morphology than uninflectes . morphological supervision is also used to augment character language models in low-resource languages .
Approach: They add morphological supervision to character language models via multitasking to improve BPC performance across 24 languages even when morphology data and language modeling data are disjointed.
Outcome: The addition improves performance even when morphology data and language modeling data are disjointed.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations