Papers by Jón Guðnason

5 papers
Language Technology Programme for Icelandic 2019-2023 (2020.lrec-1)

Copied to clipboard

Challenge: a new national language technology programme for Icelandic is described . the programme aims to make Icelandic usable in communication and interactions in the digital world .
Approach: They describe a new national language technology programme for Icelandic . the programme aims to make Icelandic usable in communication and interactions in the digital world .
Outcome: The proposed programme aims to make Icelandic usable in communication and interactions in the digital world.
Samrómur: Crowd-sourcing large amounts of data (2022.lrec-1)

Copied to clipboard

Challenge: Samrómur is the largest prompted speech collection effort for Icelandic so far and verification is as monumental as the collection itself.
Approach: They propose to collect large and diverse corpus for automatic speech recognition and similar tools using crowd-sourced donations.
Outcome: The collected utterances are based on the Mozilla Common Voice platform and are available for free on the Samrómur collection platform.
Risamálheild: A Very Large Icelandic Text Corpus (L18-1)

Copied to clipboard

Challenge: The corpus contains more than one billion running words from mostly contemporary texts.
Approach: They present the Icelandic Gigaword Corpus (IGC) with minimal work and resources.
Outcome: The Icelandic Gigaword Corpus (IGC) contains more than one billion running words from mostly contemporary texts.
Samrómur Children: An Icelandic Speech Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Samrómur Children contains 131 hours of read speech from Icelandic children aged between 4 to 17 years.
Approach: They propose to build a large-scale speech corpus for automatic speech recognition for Icelandic.
Outcome: The corpus contains 131 hours of read speech from Icelandic children aged 4 to 17 years . the goal of the project is to make Icelandic available in language-technology applications .
Open ASR for Icelandic: Resources and a Baseline System (L18-1)

Copied to clipboard

Challenge: Existing language resources are not sufficient for less-resourced languages, but a system with sufficient resources is needed.
Approach: They describe available language resources and their preparation for use in a large vocabulary speech recognition system for Icelandic.
Outcome: The proposed system improves on acoustic training sets and a speech corpus with a pronunciation dictionary.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations