Papers by Hafsteinn Einarsson

6 papers
Gendered Grammar or Ingrained Bias? Exploring Gender Bias in Icelandic Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Large language models are trained on vast datasets and exhibit increased output quality in proportion to the amount of data that is used to train them.
Approach: They explore whether language models mirror gender distributions within professions or exhibit biases tied to their grammatical genders.
Outcome: The proposed model may reflect and amplify gender bias, racism, religious prejudice, and queerphobia in training data that may not always be recent.
A Warm Start and a Clean Crawled Corpus - A Recipe for Good Language Models (2022.lrec-1)

Copied to clipboard

Challenge: Pre-trained neural language models have shown impressive results when adapted for a variety of classification and text generation tasks.
Approach: They propose to use Icelandic's Icelandic Common Crawl Corpus to train language models that achieve state-of-the-art performance in downstream tasks.
Outcome: The proposed models achieve state-of-the-art in a variety of downstream tasks including part-of speech tagging, named entity recognition and constituency parsing.
Good or Bad News? Exploring GPT-4 for Sentiment Analysis for Faroese on a Public News Corpora (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on sentiment analysis in low-resource languages have focused on major languages and emotionally laden text genres like social media and reviews.
Approach: They propose to use GPT-4 for sentiment analysis on Faroese news texts using a multi-class approach with 225 sentences analysed in 170 articles.
Outcome: The proposed model performs remarkably well on 225 sentences and 170 articles compared to human annotators .
Applications of BERT Models Towards Automation of Clinical Coding in Icelandic (2024.findings-naacl)

Copied to clipboard

Challenge: Traditionally, clinical coding is manual and laborintensive task prone to human error.
Approach: They analyze 25 years of electronic health records from the Landspitali University Hospital in Icelandic to explore the potential of using NLP for clinical coding.
Outcome: The best-performing model achieves competitive results in micro and macro F1 scores, with label attention contributing significantly to its success.
GameQA: Gamified Mobile App Platform for Building Multiple-Domain Question-Answering Datasets (2023.eacl-demo)

Copied to clipboard

Challenge: a common problem with question-answering datasets is that they require annotators to source answers from the internet . a crowd-sourcing platform is available for low-resource languages, but it is limited in terms of information available.
Approach: They propose a crowd-sourcing platform to gather multiple-domain QA data for low-resource languages.
Outcome: The proposed platform rivals large QA datasets for high-resource languages in size and answerability.
Natural Questions in Icelandic (2022.lrec-1)

Copied to clipboard

Challenge: Developing such datasets is important for the development and evaluation of Icelandic QA systems.
Approach: They present the first extractive question answering dataset for Icelandic, Natural Questions in Icelandic.
Outcome: The proposed dataset is a valuable resource for Icelandic which is being evaluated by a team of researchers.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations