Papers by Vilhjalmur Thorsteinsson
A Warm Start and a Clean Crawled Corpus - A Recipe for Good Language Models (2022.lrec-1)
Copied to clipboard
Vésteinn Snæbjarnarson, Haukur Barri Símonarson, Pétur Orri Ragnarsson, Svanhvít Lilja Ingólfsdóttir, Haukur Jónsson, Vilhjalmur Thorsteinsson, Hafsteinn Einarsson
| Challenge: | Pre-trained neural language models have shown impressive results when adapted for a variety of classification and text generation tasks. |
| Approach: | They propose to use Icelandic's Icelandic Common Crawl Corpus to train language models that achieve state-of-the-art performance in downstream tasks. |
| Outcome: | The proposed models achieve state-of-the-art in a variety of downstream tasks including part-of speech tagging, named entity recognition and constituency parsing. |
Byte-Level Grammatical Error Correction Using Synthetic and Curated Corpora (2023.acl-long)
Copied to clipboard
Svanhvít Lilja Ingólfsdóttir, Petur Ragnarsson, Haukur Jónsson, Haukur Simonarson, Vilhjalmur Thorsteinsson, Vésteinn Snæbjarnarson
| Challenge: | Spelling mistakes due to typos and rushed writing, nonstandard punctuation and spelling, and grammatical and stylistic issues are common to almost everyone who writes any kind of text. |
| Approach: | They propose to use a common subword unit vocabulary and byte-level encoding to fine tune two subword-level models and one byte level model on hand-corrected error corpora. |
| Outcome: | The proposed model improves accuracy for spelling and grammatical errors and more complex errors. |