Papers by Hafsteinn Einarsson
Gendered Grammar or Ingrained Bias? Exploring Gender Bias in Icelandic Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Large language models are trained on vast datasets and exhibit increased output quality in proportion to the amount of data that is used to train them. |
| Approach: | They explore whether language models mirror gender distributions within professions or exhibit biases tied to their grammatical genders. |
| Outcome: | The proposed model may reflect and amplify gender bias, racism, religious prejudice, and queerphobia in training data that may not always be recent. |
A Warm Start and a Clean Crawled Corpus - A Recipe for Good Language Models (2022.lrec-1)
Copied to clipboard
Vésteinn Snæbjarnarson, Haukur Barri Símonarson, Pétur Orri Ragnarsson, Svanhvít Lilja Ingólfsdóttir, Haukur Jónsson, Vilhjalmur Thorsteinsson, Hafsteinn Einarsson
| Challenge: | Pre-trained neural language models have shown impressive results when adapted for a variety of classification and text generation tasks. |
| Approach: | They propose to use Icelandic's Icelandic Common Crawl Corpus to train language models that achieve state-of-the-art performance in downstream tasks. |
| Outcome: | The proposed models achieve state-of-the-art in a variety of downstream tasks including part-of speech tagging, named entity recognition and constituency parsing. |
Good or Bad News? Exploring GPT-4 for Sentiment Analysis for Faroese on a Public News Corpora (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies on sentiment analysis in low-resource languages have focused on major languages and emotionally laden text genres like social media and reviews. |
| Approach: | They propose to use GPT-4 for sentiment analysis on Faroese news texts using a multi-class approach with 225 sentences analysed in 170 articles. |
| Outcome: | The proposed model performs remarkably well on 225 sentences and 170 articles compared to human annotators . |
Applications of BERT Models Towards Automation of Clinical Coding in Icelandic (2024.findings-naacl)
Copied to clipboard
| Challenge: | Traditionally, clinical coding is manual and laborintensive task prone to human error. |
| Approach: | They analyze 25 years of electronic health records from the Landspitali University Hospital in Icelandic to explore the potential of using NLP for clinical coding. |
| Outcome: | The best-performing model achieves competitive results in micro and macro F1 scores, with label attention contributing significantly to its success. |
GameQA: Gamified Mobile App Platform for Building Multiple-Domain Question-Answering Datasets (2023.eacl-demo)
Copied to clipboard
Njall Skarphedinsson, Breki Gudmundsson, Steinar Smari, Marta Kristin Larusdottir, Hafsteinn Einarsson, Abuzar Khan, Eric Nyberg, Hrafn Loftsson
| Challenge: | a common problem with question-answering datasets is that they require annotators to source answers from the internet . a crowd-sourcing platform is available for low-resource languages, but it is limited in terms of information available. |
| Approach: | They propose a crowd-sourcing platform to gather multiple-domain QA data for low-resource languages. |
| Outcome: | The proposed platform rivals large QA datasets for high-resource languages in size and answerability. |
Natural Questions in Icelandic (2022.lrec-1)
Copied to clipboard
| Challenge: | Developing such datasets is important for the development and evaluation of Icelandic QA systems. |
| Approach: | They present the first extractive question answering dataset for Icelandic, Natural Questions in Icelandic. |
| Outcome: | The proposed dataset is a valuable resource for Icelandic which is being evaluated by a team of researchers. |