Lilja Øvrelid, Andre Kåsen, Kristin Hagen, Anders Nøklestad, Per Erik Solberg, Janne Bondi Johannessen
| Challenge: | a long-term goal of this work is to develop a parser for spoken Norwegian with the immediate goal of parsing the whole LIA material. |
| Approach: | They describe the LIA treebank of transcribed spoken Norwegian dialects and their transcription, transliteration and further morphosyntactic annotation. |
| Outcome: | The treebank consists of 13,608 tokens, distributed over 1396 segments taken from three different dialects of spoken Norwegian. |
Similar Papers
The Norwegian Dialect Corpus Treebank (2022.lrec-1)
Copied to clipboard
Andre Kåsen, Kristin Hagen, Anders Nøklestad, Joel Priestly, Per Erik Solberg, Dag Trygve Truslew Haug
| Challenge: | The NDC Treebank consists of recordings made between 2006 and 2012 and is annotated with morphological and syntactic information. |
| Approach: | They present the NDC Treebank of spoken Norwegian dialects in the Bokml variety of Norwegian. |
| Outcome: | The treebank consists of 4587 speech segments, overall 66009 tokens, from 17 different Norwegian dialects from south, west, east and north of Norway. |
Spoken Language Treebanks in Universal Dependencies: an Overview (2022.lrec-1)
Copied to clipboard
| Challenge: | spoken language treebanks have divergent annotation schemes limiting cross-resource explorations . many spoken language trees have no written form, but many of the world languages have no spoken form at all. |
| Approach: | They propose to use the Universal Dependencies annotation scheme to annotate spoken language treebanks using a morphosyntactic annotation scheme. |
| Outcome: | The proposed treebanks differ significantly with respect to the inventory and format of transcribed phenomena and the principles adopted in their morphosyntactic annotation. |
A Universal Dependencies Treebank of Ancient Hebrew (2022.lrec-1)
Copied to clipboard
| Challenge: | Using a rule-based parser, we construct a treebank with morphological annotations of Ancient Hebrew . the Hebrew Scriptures are a collection of 39 books written in the first millennium BC in Ancient Hebrew. |
| Approach: | They propose to use a Universal Dependencies treebank with morphological annotations of Ancient Hebrew for comparative study with ancient translations and analysis of Hebrew syntax. |
| Outcome: | The proposed treebank can be used in comparative study with ancient translations and analysis of Hebrew syntax. |
Aligning the Norwegian UD Treebank with Entity and Coreference Information (2024.lrec-main)
Copied to clipboard
| Challenge: | merged corpora of entity and coreference data are presented for the two written forms of Norwegian: Bokml and Nynorsk. |
| Approach: | They propose to combine entity and coreference data from two UD treebanks for Norwegian written forms: Bokml and Nynorsk. |
| Outcome: | The merged corpora comprise the first Norwegian UD treebank enriched with named entities and coreference information, supporting the standardized format for the CorefUD initiative. |
NorNE: Annotating Named Entities for Norwegian (2020.lrec-1)
Copied to clipboard
| Challenge: | Using the annotations of the existing treebank, we have created a dataset for named entity recognition for Norwegian. |
| Approach: | They propose to create a manually annotated corpus of named entities for Norwegian . they propose to add named entity annotations to existing treebank . |
| Outcome: | The proposed dataset extends the annotation of the existing Norwegian Dependency Treebank. |
Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection (2020.lrec-1)
Copied to clipboard
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, Daniel Zeman
| Challenge: | Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages. |
| Approach: | They describe version 2 of the universal guidelines and discuss major changes from UD v1 to UD 2 . they propose a morphological layer, a syntactic layer and a word segmentation layer . |
| Outcome: | The proposed treebanks are available for 90 languages and have been updated to meet the needs of multilingual parsers and researchers. |
NB Uttale: A Norwegian Pronunciation Lexicon with Dialect Variation (2024.lrec-main)
Copied to clipboard
| Challenge: | lexicon is based on the NST Bokml lexiconic for East Norwegian . lexica are an essential linguistic resource in speech recognition and speech synthesis systems . |
| Approach: | They propose to use Bokml orthographic word forms and up to eight alternate phonological transcriptions per word form to generate a Norwegian pronunciation lexicon. |
| Outcome: | The proposed model improves the accuracy of the proposed model and its outputs with word- and phoneme-error-rate metrics. |
Corpus Creation and Automatic Alignment of Historical Dutch Dialect Speech (2024.lrec-main)
Copied to clipboard
| Challenge: | The Dutch Dialect Database contains dialectal variations of Dutch recorded in the second half of the twentieth century. |
| Approach: | They propose to create a corpus containing audio recordings and orthographic transcriptions of Dutch dialects recorded in the second half of the 20th century. |
| Outcome: | The Dutch Dialect Database contains dialectal variations recorded all over the Netherlands in the second half of the twentieth century. |
The Norwegian Parliamentary Speech Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | the dataset contains recordings of meetings at the Norwegian parliament . it is the first publicly available dataset containing unscripted, Norwegian speech . |
| Approach: | the Norwegian Parliamentary Speech Corpus is a publicly available speech dataset . it contains recordings of meetings from the Norwegian parliament with orthographic transcriptions . the dataset is intended to fill a gap in the available unscripted speech data . |
| Outcome: | the dataset contains recordings of meetings at the Norwegian parliament with orthographic transcriptions in Norwegian Bokml and Norwegian Nynorsk. |
Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell (2020.acl-main)
Copied to clipboard
Djamé Seddah, Farah Essaidi, Amal Fethi, Matthieu Futeral, Benjamin Muller, Pedro Javier Ortiz Suárez, Benoît Sagot, Abhishek Srivastava
| Challenge: | a treebank for a north-African Arabic dialect known for code-switching is made freely available . authors: geopolitical events are a factor highlighting a language deficiency in terms of natural language processing resources . |
| Approach: | They propose to make a treebank for a romanized user-generated content variety of Algerian . they supplement it with 50k unlabeled sentences from common crawl and web-crawled data . |
| Outcome: | The proposed treebank is made of 1500 sentences, fully annotated in morpho-syntax and universal dependency syntax, with full translation at both the word and sentence levels. |