The LIA Treebank of Spoken Norwegian Dialects (L18-1)

Copied to clipboard

Challenge: a long-term goal of this work is to develop a parser for spoken Norwegian with the immediate goal of parsing the whole LIA material.
Approach: They describe the LIA treebank of transcribed spoken Norwegian dialects and their transcription, transliteration and further morphosyntactic annotation.
Outcome: The treebank consists of 13,608 tokens, distributed over 1396 segments taken from three different dialects of spoken Norwegian.

Similar Papers

The Norwegian Dialect Corpus Treebank (2022.lrec-1)

Copied to clipboard

Challenge: The NDC Treebank consists of recordings made between 2006 and 2012 and is annotated with morphological and syntactic information.
Approach: They present the NDC Treebank of spoken Norwegian dialects in the Bokml variety of Norwegian.
Outcome: The treebank consists of 4587 speech segments, overall 66009 tokens, from 17 different Norwegian dialects from south, west, east and north of Norway.
Spoken Language Treebanks in Universal Dependencies: an Overview (2022.lrec-1)

Copied to clipboard

Challenge: spoken language treebanks have divergent annotation schemes limiting cross-resource explorations . many spoken language trees have no written form, but many of the world languages have no spoken form at all.
Approach: They propose to use the Universal Dependencies annotation scheme to annotate spoken language treebanks using a morphosyntactic annotation scheme.
Outcome: The proposed treebanks differ significantly with respect to the inventory and format of transcribed phenomena and the principles adopted in their morphosyntactic annotation.
A Universal Dependencies Treebank of Ancient Hebrew (2022.lrec-1)

Copied to clipboard

Challenge: Using a rule-based parser, we construct a treebank with morphological annotations of Ancient Hebrew . the Hebrew Scriptures are a collection of 39 books written in the first millennium BC in Ancient Hebrew.
Approach: They propose to use a Universal Dependencies treebank with morphological annotations of Ancient Hebrew for comparative study with ancient translations and analysis of Hebrew syntax.
Outcome: The proposed treebank can be used in comparative study with ancient translations and analysis of Hebrew syntax.
Aligning the Norwegian UD Treebank with Entity and Coreference Information (2024.lrec-main)

Copied to clipboard

Challenge: merged corpora of entity and coreference data are presented for the two written forms of Norwegian: Bokml and Nynorsk.
Approach: They propose to combine entity and coreference data from two UD treebanks for Norwegian written forms: Bokml and Nynorsk.
Outcome: The merged corpora comprise the first Norwegian UD treebank enriched with named entities and coreference information, supporting the standardized format for the CorefUD initiative.
NorNE: Annotating Named Entities for Norwegian (2020.lrec-1)

Copied to clipboard

Challenge: Using the annotations of the existing treebank, we have created a dataset for named entity recognition for Norwegian.
Approach: They propose to create a manually annotated corpus of named entities for Norwegian . they propose to add named entity annotations to existing treebank .
Outcome: The proposed dataset extends the annotation of the existing Norwegian Dependency Treebank.
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)

Copied to clipboard

Challenge: Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages.
Approach: They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer .
Outcome: The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers.
NB Uttale: A Norwegian Pronunciation Lexicon with Dialect Variation (2024.lrec-main)

Copied to clipboard

Challenge: lexicon is based on the NST Bokml lexiconic for East Norwegian . lexica are an essential linguistic resource in speech recognition and speech synthesis systems .
Approach: They propose to use Bokml orthographic word forms and up to eight alternate phonological transcriptions per word form to generate a Norwegian pronunciation lexicon.
Outcome: The proposed model improves the accuracy of the proposed model and its outputs with word- and phoneme-error-rate metrics.
Corpus Creation and Automatic Alignment of Historical Dutch Dialect Speech (2024.lrec-main)

Copied to clipboard

Challenge: The Dutch Dialect Database contains dialectal variations of Dutch recorded in the second half of the twentieth century.
Approach: They propose to create a corpus containing audio recordings and orthographic transcriptions of Dutch dialects recorded in the second half of the 20th century.
Outcome: The Dutch Dialect Database contains dialectal variations recorded all over the Netherlands in the second half of the twentieth century.
The Norwegian Parliamentary Speech Corpus (2022.lrec-1)

Copied to clipboard

Challenge: the dataset contains recordings of meetings at the Norwegian parliament . it is the first publicly available dataset containing unscripted, Norwegian speech .
Approach: the Norwegian Parliamentary Speech Corpus is a publicly available speech dataset . it contains recordings of meetings from the Norwegian parliament with orthographic transcriptions . the dataset is intended to fill a gap in the available unscripted speech data .
Outcome: the dataset contains recordings of meetings at the Norwegian parliament with orthographic transcriptions in Norwegian Bokml and Norwegian Nynorsk.
Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell (2020.acl-main)

Copied to clipboard

Challenge: a treebank for a north-African Arabic dialect known for code-switching is made freely available . authors: geopolitical events are a factor highlighting a language deficiency in terms of natural language processing resources .
Approach: They propose to make a treebank for a romanized user-generated content variety of Algerian . they supplement it with 50k unlabeled sentences from common crawl and web-crawled data .
Outcome: The proposed treebank is made of 1500 sentences, fully annotated in morpho-syntax and universal dependency syntax, with full translation at both the word and sentence levels.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations