Papers by Vilhjalmur Thorsteinsson

2 papers
A Warm Start and a Clean Crawled Corpus - A Recipe for Good Language Models (2022.lrec-1)

Copied to clipboard

Challenge: Pre-trained neural language models have shown impressive results when adapted for a variety of classification and text generation tasks.
Approach: They propose to use Icelandic's Icelandic Common Crawl Corpus to train language models that achieve state-of-the-art performance in downstream tasks.
Outcome: The proposed models achieve state-of-the-art in a variety of downstream tasks including part-of speech tagging, named entity recognition and constituency parsing.
Byte-Level Grammatical Error Correction Using Synthetic and Curated Corpora (2023.acl-long)

Copied to clipboard

Challenge: Spelling mistakes due to typos and rushed writing, nonstandard punctuation and spelling, and grammatical and stylistic issues are common to almost everyone who writes any kind of text.
Approach: They propose to use a common subword unit vocabulary and byte-level encoding to fine tune two subword-level models and one byte level model on hand-corrected error corpora.
Outcome: The proposed model improves accuracy for spelling and grammatical errors and more complex errors.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations