| Challenge: | Eesthetic is a comprehensive Estonian noun and verb lexicon . it documents 5475 nouns inflecting for 28 paradigm cells and 5076 verbs inflection for 51 cells. |
| Approach: | They propose to use Ekilex to generate an Estonian noun and verb lexicon with a set of rules for automatic transcription. |
| Outcome: | The Estonian lexicon is based on the Ekilex database and is openly accessible . it contains a total of 452885 inflected forms and is structured and formatted as a set of CSV tables linked by formal relationships. |
Similar Papers
A multilingual collection of CoNLL-U-compatible morphological lexicons (L18-1)
Copied to clipboard
| Challenge: | Existing morphological lexicons are limited in scope and are not universally accepted . morphology lexical information is encoded into morphologists or gathered in lexiconics . |
| Approach: | They propose a multilingual collection of morphological lexicons that follow the Universal Dependencies initiative. |
| Outcome: | The proposed collection of 53 morphological lexicons covers 38 languages . they have been shown to improve part-of-speech tagging and parsing accuracy . |
Unifying Morphology Resources with OntoLex-Morph. A Case Study in German (2022.lrec-1)
Copied to clipboard
| Challenge: | OntoLex is a widely used community standard for machine-readable lexical resources on the web. |
| Approach: | They propose a module for representing morphology that can be used to encode and integrate morphological resources on a unified basis. |
| Outcome: | The proposed module can be used to represent morphological resources on a unified basis. |
Opening the Romance Verbal Inflection Dataset 2.0: A CLDF lexicon (2020.lrec-1)
Copied to clipboard
| Challenge: | lexicon provides verbal paradigm forms in broad IPA phonemic notation for 74 varieties . most resources used to study language evolution computationally rely on multilingual contemporary information . |
| Approach: | They propose a multilingual lexicon of Romance inflection covering 74 varieties . they annotate verbal paradigm forms in broad IPA phonemic notation and organize paradigm cells to reflect cognacy . |
| Outcome: | The lexicon provides verbal paradigm forms in broad IPA phonemic notation for 74 varieties. |
Morphological Reinflection with Multiple Arguments: An Extended Annotation schema and a Georgian Case Study (2022.acl-short)
Copied to clipboard
| Challenge: | morphological annotations are a common problem in some languages, but the flat structure of the current schema makes it impossible to treat them. |
| Approach: | They propose a general solution for polypersonal agreement in Georgian language . they extend the existing UniMorph annotation schema to address this problem . |
| Outcome: | The proposed framework covers all possible variants of argument marking, and is accurate and balanced. |
Wikinflection Corpus: A (Better) Multilingual, Morpheme-Annotated Inflectional Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Inflectional corpora with annotated morpheme boundaries are scarce in the NLP community . a generated, multilingual inflectional lexicon with morphological features is not as good as UniMorph's . |
| Approach: | They evaluate a multilingual inflectional corpus with morpheme boundaries from the English Wiktionary and the UniMorph project's inflection corpus. |
| Outcome: | The generated Wikinflection corpus is not as good as UniMorph's, but extracts significant amount of words from the intersection of the two corpora. |
Unicode Normalization and Grapheme Parsing of Indic Languages (2024.lrec-main)
Copied to clipboard
Nazmuddoha Ansary, Quazi Adibur Rahman Adib, Tahsin Reasat, Asif Shahriyar Sushmit, Ahmed Imtiaz Humayun, Sazia Mehnaz, Kanij Fatema, Mohammad Mamun Or Rashid, Farig Sadeque
| Challenge: | Indic writing systems encode words as linear sequences of Unicode characters . authors propose a grapheme parser for Abugida text to normalize inconsistencies . |
| Approach: | They propose a normalizer for normalizing inconsistencies caused by Unicode encoding schemes . grapheme parser for Abugida deconstructs words into visually distinct orthographic syllables . |
| Outcome: | The proposed library is more efficient than the previously used IndicNLP normalizer . it deconstructs words into visually distinct orthographic syllables or complex graphemes . |
LinguaMeta: Unified Metadata for Thousands of Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | LinguaMeta is a unified repository of language metadata for thousands of languages. |
| Approach: | They introduce LinguaMeta, a unified resource for language metadata for thousands of languages. |
| Outcome: | The proposed resource is intended for use by researchers and organizations who aim to extend technology to thousands of languages. |
A Multi-word Expression Dataset for Swedish (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing data on compositionality of multi-word expressions is limited and only available for high resource languages. |
| Approach: | They present a set of Swedish multi-word expressions annotated with degree of compositionality . they also consider syntactically complex constructions and publish a formal specification of each expression . |
| Outcome: | The proposed dataset includes 96 Swedish multi-word expressions with degree of compositionality. |
GeCzLex: Lexicon of Czech and German Anaphoric Connectives (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing lexicons of connectives are interlinked with each other to provide a bilingual inventory of connective entries. |
| Approach: | They introduce the first version of a lexicon for translation equivalents of Czech and German discourse connectives. |
| Outcome: | The lexicon is the first bilingual inventory of connectives with linkage on the level of individual entries. |
Towards a Semi-Automatic Detection of Reflexive and Reciprocal Constructions and Their Representation in a Valency Lexicon (2020.lrec-1)
Copied to clipboard
| Challenge: | valency lexicons describe valencies of verbs in non-reflexive and non-reciprocal constructions . reflexive and reciprocal constructions are common morphosyntactic forms of verb . |
| Approach: | They propose a semi-automatic procedure to detect verbs with reflexive and reciprocal constructions in corpus data. |
| Outcome: | The proposed procedure detects verbs that form reflexive and reciprocal constructions in corpus data. |