Eesthetic: A Paralex Lexicon of Estonian Paradigms (2024.lrec-main)

Copied to clipboard

Challenge: Eesthetic is a comprehensive Estonian noun and verb lexicon . it documents 5475 nouns inflecting for 28 paradigm cells and 5076 verbs inflection for 51 cells.
Approach: They propose to use Ekilex to generate an Estonian noun and verb lexicon with a set of rules for automatic transcription.
Outcome: The Estonian lexicon is based on the Ekilex database and is openly accessible . it contains a total of 452885 inflected forms and is structured and formatted as a set of CSV tables linked by formal relationships.

Similar Papers

A multilingual collection of CoNLL-U-compatible morphological lexicons (L18-1)

Copied to clipboard

Challenge: Existing morphological lexicons are limited in scope and are not universally accepted . morphology lexical information is encoded into morphologists or gathered in lexiconics .
Approach: They propose a multilingual collection of morphological lexicons that follow the Universal Dependencies initiative.
Outcome: The proposed collection of 53 morphological lexicons covers 38 languages . they have been shown to improve part-of-speech tagging and parsing accuracy .
Unifying Morphology Resources with OntoLex-Morph. A Case Study in German (2022.lrec-1)

Copied to clipboard

Challenge: OntoLex is a widely used community standard for machine-readable lexical resources on the web.
Approach: They propose a module for representing morphology that can be used to encode and integrate morphological resources on a unified basis.
Outcome: The proposed module can be used to represent morphological resources on a unified basis.
Opening the Romance Verbal Inflection Dataset 2.0: A CLDF lexicon (2020.lrec-1)

Copied to clipboard

Challenge: lexicon provides verbal paradigm forms in broad IPA phonemic notation for 74 varieties . most resources used to study language evolution computationally rely on multilingual contemporary information .
Approach: They propose a multilingual lexicon of Romance inflection covering 74 varieties . they annotate verbal paradigm forms in broad IPA phonemic notation and organize paradigm cells to reflect cognacy .
Outcome: The lexicon provides verbal paradigm forms in broad IPA phonemic notation for 74 varieties.
Morphological Reinflection with Multiple Arguments: An Extended Annotation schema and a Georgian Case Study (2022.acl-short)

Copied to clipboard

Challenge: morphological annotations are a common problem in some languages, but the flat structure of the current schema makes it impossible to treat them.
Approach: They propose a general solution for polypersonal agreement in Georgian language . they extend the existing UniMorph annotation schema to address this problem .
Outcome: The proposed framework covers all possible variants of argument marking, and is accurate and balanced.
Wikinflection Corpus: A (Better) Multilingual, Morpheme-Annotated Inflectional Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Inflectional corpora with annotated morpheme boundaries are scarce in the NLP community . a generated, multilingual inflectional lexicon with morphological features is not as good as UniMorph's .
Approach: They evaluate a multilingual inflectional corpus with morpheme boundaries from the English Wiktionary and the UniMorph project's inflection corpus.
Outcome: The generated Wikinflection corpus is not as good as UniMorph's, but extracts significant amount of words from the intersection of the two corpora.
Unicode Normalization and Grapheme Parsing of Indic Languages (2024.lrec-main)

Copied to clipboard

Challenge: Indic writing systems encode words as linear sequences of Unicode characters . authors propose a grapheme parser for Abugida text to normalize inconsistencies .
Approach: They propose a normalizer for normalizing inconsistencies caused by Unicode encoding schemes . grapheme parser for Abugida deconstructs words into visually distinct orthographic syllables .
Outcome: The proposed library is more efficient than the previously used IndicNLP normalizer . it deconstructs words into visually distinct orthographic syllables or complex graphemes .
LinguaMeta: Unified Metadata for Thousands of Languages (2024.lrec-main)

Copied to clipboard

Challenge: LinguaMeta is a unified repository of language metadata for thousands of languages.
Approach: They introduce LinguaMeta, a unified resource for language metadata for thousands of languages.
Outcome: The proposed resource is intended for use by researchers and organizations who aim to extend technology to thousands of languages.
A Multi-word Expression Dataset for Swedish (2020.lrec-1)

Copied to clipboard

Challenge: Existing data on compositionality of multi-word expressions is limited and only available for high resource languages.
Approach: They present a set of Swedish multi-word expressions annotated with degree of compositionality . they also consider syntactically complex constructions and publish a formal specification of each expression .
Outcome: The proposed dataset includes 96 Swedish multi-word expressions with degree of compositionality.
GeCzLex: Lexicon of Czech and German Anaphoric Connectives (2020.lrec-1)

Copied to clipboard

Challenge: Existing lexicons of connectives are interlinked with each other to provide a bilingual inventory of connective entries.
Approach: They introduce the first version of a lexicon for translation equivalents of Czech and German discourse connectives.
Outcome: The lexicon is the first bilingual inventory of connectives with linkage on the level of individual entries.
Towards a Semi-Automatic Detection of Reflexive and Reciprocal Constructions and Their Representation in a Valency Lexicon (2020.lrec-1)

Copied to clipboard

Challenge: valency lexicons describe valencies of verbs in non-reflexive and non-reciprocal constructions . reflexive and reciprocal constructions are common morphosyntactic forms of verb .
Approach: They propose a semi-automatic procedure to detect verbs with reflexive and reciprocal constructions in corpus data.
Outcome: The proposed procedure detects verbs that form reflexive and reciprocal constructions in corpus data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations