Papers by Anders Nøklestad

3 papers
The Norwegian Dialect Corpus Treebank (2022.lrec-1)

Copied to clipboard

Challenge: The NDC Treebank consists of recordings made between 2006 and 2012 and is annotated with morphological and syntactic information.
Approach: They present the NDC Treebank of spoken Norwegian dialects in the Bokml variety of Norwegian.
Outcome: The treebank consists of 4587 speech segments, overall 66009 tokens, from 17 different Norwegian dialects from south, west, east and north of Norway.
The LIA Treebank of Spoken Norwegian Dialects (L18-1)

Copied to clipboard

Challenge: a long-term goal of this work is to develop a parser for spoken Norwegian with the immediate goal of parsing the whole LIA material.
Approach: They describe the LIA treebank of transcribed spoken Norwegian dialects and their transcription, transliteration and further morphosyntactic annotation.
Outcome: The treebank consists of 13,608 tokens, distributed over 1396 segments taken from three different dialects of spoken Norwegian.
Comparing Methods for Measuring Dialect Similarity in Norwegian (2020.lrec-1)

Copied to clipboard

Challenge: a coarse-grained transcription of speech is sufficient to replicate dialectal boundaries, but it can be generalised over by an automatic method.
Approach: They propose to use two different methods to measure dialect similarity in Norwegian . they use the Levenshtein method and the neural long short term memory algorithm . the paper shows that coarse-grained transcriptions of speech can generate dialect maps .
Outcome: The proposed method can generalise over coarse-grained transcriptions, but it needs a large dataset . the proposed method is compared with canonical maps found in the dialect literature .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations