Papers by Anders Nøklestad
The Norwegian Dialect Corpus Treebank (2022.lrec-1)
Copied to clipboard
Andre Kåsen, Kristin Hagen, Anders Nøklestad, Joel Priestly, Per Erik Solberg, Dag Trygve Truslew Haug
| Challenge: | The NDC Treebank consists of recordings made between 2006 and 2012 and is annotated with morphological and syntactic information. |
| Approach: | They present the NDC Treebank of spoken Norwegian dialects in the Bokml variety of Norwegian. |
| Outcome: | The treebank consists of 4587 speech segments, overall 66009 tokens, from 17 different Norwegian dialects from south, west, east and north of Norway. |
The LIA Treebank of Spoken Norwegian Dialects (L18-1)
Copied to clipboard
Lilja Øvrelid, Andre Kåsen, Kristin Hagen, Anders Nøklestad, Per Erik Solberg, Janne Bondi Johannessen
| Challenge: | a long-term goal of this work is to develop a parser for spoken Norwegian with the immediate goal of parsing the whole LIA material. |
| Approach: | They describe the LIA treebank of transcribed spoken Norwegian dialects and their transcription, transliteration and further morphosyntactic annotation. |
| Outcome: | The treebank consists of 13,608 tokens, distributed over 1396 segments taken from three different dialects of spoken Norwegian. |
Comparing Methods for Measuring Dialect Similarity in Norwegian (2020.lrec-1)
Copied to clipboard
| Challenge: | a coarse-grained transcription of speech is sufficient to replicate dialectal boundaries, but it can be generalised over by an automatic method. |
| Approach: | They propose to use two different methods to measure dialect similarity in Norwegian . they use the Levenshtein method and the neural long short term memory algorithm . the paper shows that coarse-grained transcriptions of speech can generate dialect maps . |
| Outcome: | The proposed method can generalise over coarse-grained transcriptions, but it needs a large dataset . the proposed method is compared with canonical maps found in the dialect literature . |