Universal Morphologies for the Caucasus region (L18-1)

Copied to clipboard

Challenge: Caucasus region is famed for its rich and diverse arrays of languages and language families . authors describe efforts to improve the coverage of Universal Morphologies for languages of the region .
Approach: They propose to improve the coverage of Universal Morphologies for Caucasus languages . they propose to complement the Universal Dependencies which focus on morphosyntax and syntax.
Outcome: The proposed framework improves the coverage of languages of the Caucasus region . the proposed framework criticizes the UniMorph TSV format for its limited expressiveness .

Similar Papers

UniMorph 4.0: Universal Morphology (2022.lrec-1)

Copied to clipboard

Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieraś, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Abbott Lane, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóǧa, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan North Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo M. Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Prud’hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya D. McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
Challenge: The Universal Morphology project provides broad-coverage instantiated morphological inflection tables for hundreds of diverse languages.
Approach: They propose a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema.
Outcome: The proposed schema has added 66 new languages, including 24 endangered languages.
UniMorph 3.0: Universal Morphology (2020.lrec-1)

Copied to clipboard

Challenge: Explicit modeling of morphology has demonstrable benefits for language modeling, speech recognition, word embedding and keyword search.
Approach: They propose a language-independent feature schema for rich morphological annotation and a type-level resource for annotated data in diverse languages.
Outcome: The proposed schema has been improved to make it more complete and correct, and adds 66 new languages and parts of speech for 12 languages.
UniMorph 2.0: Universal Morphology (L18-1)

Copied to clipboard

Challenge: The Universal Morphology project is a collaborative effort to improve how NLP handles complex morphology across the world's languages.
Approach: They propose to use a universal tagset to annotate morphological data using a schema that includes a lemma and a bundle of morphology features.
Outcome: The project releases annotated morphological data using a universal tagset, the UniMorph schema.
Universal Anaphora: The First Three Years (2024.lrec-main)

Copied to clipboard

Challenge: Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by expanding the aspects of anaphonic interpretation which are or can be reliably annotated in an anagraphic corpora.
Approach: They propose to develop a standard for anaphoric annotations and a method for evaluating models that can carry out this type of interpretation.
Outcome: The Universal Anaphora initiative aims to push forward the state of the art in anaphora and anaphorism resolution by producing unified standards to annotate and encode annotations, delivering datasets encoded according to these standards, and developing methods for evaluating models that carry out this type of interpretation.
The Abkhaz National Corpus (L18-1)

Copied to clipboard

Challenge: Abkhaz National Corpus is a comprehensive and open, grammatically annotated text corpus . it is currently growing and is being extended to include all important texts written in the language .
Approach: They propose to use the Abkhaz National Corpus to annotate Abkhhaz texts . the corpus is a comprehensive and open, grammatically annotated text corpus .
Outcome: The proposed corpus is a grammatically annotated text corpus which makes the language accessible to scientific investigations from various perspectives.
Morphological Reinflection with Multiple Arguments: An Extended Annotation schema and a Georgian Case Study (2022.acl-short)

Copied to clipboard

Challenge: morphological annotations are a common problem in some languages, but the flat structure of the current schema makes it impossible to treat them.
Approach: They propose a general solution for polypersonal agreement in Georgian language . they extend the existing UniMorph annotation schema to address this problem .
Outcome: The proposed framework covers all possible variants of argument marking, and is accurate and balanced.
Wikinflection Corpus: A (Better) Multilingual, Morpheme-Annotated Inflectional Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Inflectional corpora with annotated morpheme boundaries are scarce in the NLP community . a generated, multilingual inflectional lexicon with morphological features is not as good as UniMorph's .
Approach: They evaluate a multilingual inflectional corpus with morpheme boundaries from the English Wiktionary and the UniMorph project's inflection corpus.
Outcome: The generated Wikinflection corpus is not as good as UniMorph's, but extracts significant amount of words from the intersection of the two corpora.
K-UniMorph: Korean Universal Morphology and its Feature Schema (2023.findings-acl)

Copied to clipboard

Challenge: Previously, the Korean language has been underrepresented in the field of morphological paradigms amongst hundreds of diverse world languages.
Approach: They propose a new Universal Morphology dataset for Korean that preserves its distinct characteristics.
Outcome: The proposed dataset extracts inflected Korean verb forms from the largest annotated corpus for Korean.
Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages (2020.lrec-1)

Copied to clipboard

Challenge: Using morphosyntactic tools, we train and distribute tools for approximately one thousand languages.
Approach: They train and distribute morphosyntactic tools for approximately one thousand languages.
Outcome: The results show that the tools generalize well across rare and common forms alike.
LinguaMeta: Unified Metadata for Thousands of Languages (2024.lrec-main)

Copied to clipboard

Challenge: LinguaMeta is a unified repository of language metadata for thousands of languages.
Approach: They introduce LinguaMeta, a unified resource for language metadata for thousands of languages.
Outcome: The proposed resource is intended for use by researchers and organizations who aim to extend technology to thousands of languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations