Papers by Petter Mæhlum

5 papers
A Fine-grained Sentiment Dataset for Norwegian (2020.lrec-1)

Copied to clipboard

Challenge: Using a dataset for fine-grained sentiment analysis in Norwegian, we examine the annotation effort and provide an overview of the developed annotation guidelines.
Approach: They propose a dataset for fine-grained sentiment analysis in Norwegian . they provide an overview of the developed annotation guidelines and analyze inter-annotator agreement .
Outcome: The proposed dataset is the first of its kind for Norwegian and is available online.
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT) (2025.acl-long)

Copied to clipboard

Challenge: a large number of textual data is needed to train state-of-the-art large language models.
Approach: They propose a collection of monolingual and parallel corpora from the Internet Archive . they document the entire data pipeline and release the code to reproduce it .
Outcome: The proposed collection of monolingual and parallel corpora is based on the HPLT v2 dataset . it includes 8T tokens covering 193 languages and 380M sentence pairs covering 51 languages .
Estimating Lexical Complexity from Document-Level Distributions (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for complexity estimation are limited to entire documents . health assessment tools are too short for existing methods to apply .
Approach: They propose a two-step approach for estimating lexical complexity that does not rely on pre-annotated data.
Outcome: The proposed method is tested on the Norwegian language and compares with other assessment tools.
NorDiaChange: Diachronic Semantic Change Dataset for Norwegian (2022.lrec-1)

Copied to clipboard

Challenge: NorDiaChange is the first dataset of diachronic semantic change on the lexical level for Norwegian.
Approach: They describe a manual annotation process for a new dataset of diachronic semantic change for Norwegian.
Outcome: The proposed dataset covers the time periods related to pre- and post-war events, oil and gas discovery in Norway, and technological developments.
EDEN: A Dataset for Event Detection in Norwegian News (2024.lrec-main)

Copied to clipboard

Challenge: EDEN is the first dataset annotated with event information at the sentence level for the Norwegian language.
Approach: They propose to annotate Norwegian news text and transcribed speech using ACE event schema.
Outcome: The proposed dataset is the first annotated dataset for Norwegian, with a language-specific annotation process.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations