Papers by Petter Mæhlum
A Fine-grained Sentiment Dataset for Norwegian (2020.lrec-1)
Copied to clipboard
| Challenge: | Using a dataset for fine-grained sentiment analysis in Norwegian, we examine the annotation effort and provide an overview of the developed annotation guidelines. |
| Approach: | They propose a dataset for fine-grained sentiment analysis in Norwegian . they provide an overview of the developed annotation guidelines and analyze inter-annotator agreement . |
| Outcome: | The proposed dataset is the first of its kind for Norwegian and is available online. |
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT) (2025.acl-long)
Copied to clipboard
Laurie Burchell, Ona De Gibert Bonet, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Pinzhen Chen, Mariia Fedorova, Liane Guillou, Barry Haddow, Jan Hajič, Jindřich Helcl, Erik Henriksson, Mateusz Klimaszewski, Ville Komulainen, Andrey Kutuzov, Joona Kytöniemi, Veronika Laippala, Petter Mæhlum, Bhavitvya Malik, Farrokh Mehryary, Vladislav Mikhailov, Nikita Moghe, Amanda Myntti, Dayyán O’Brien, Stephan Oepen, Proyag Pal, Jousia Piha, Sampo Pyysalo, Gema Ramírez-Sánchez, David Samuel, Pavel Stepachev, Jörg Tiedemann, Dušan Variš, Tereza Vojtěchová, Jaume Zaragoza-Bernabeu
| Challenge: | a large number of textual data is needed to train state-of-the-art large language models. |
| Approach: | They propose a collection of monolingual and parallel corpora from the Internet Archive . they document the entire data pipeline and release the code to reproduce it . |
| Outcome: | The proposed collection of monolingual and parallel corpora is based on the HPLT v2 dataset . it includes 8T tokens covering 193 languages and 380M sentence pairs covering 51 languages . |
Estimating Lexical Complexity from Document-Level Distributions (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for complexity estimation are limited to entire documents . health assessment tools are too short for existing methods to apply . |
| Approach: | They propose a two-step approach for estimating lexical complexity that does not rely on pre-annotated data. |
| Outcome: | The proposed method is tested on the Norwegian language and compares with other assessment tools. |
NorDiaChange: Diachronic Semantic Change Dataset for Norwegian (2022.lrec-1)
Copied to clipboard
| Challenge: | NorDiaChange is the first dataset of diachronic semantic change on the lexical level for Norwegian. |
| Approach: | They describe a manual annotation process for a new dataset of diachronic semantic change for Norwegian. |
| Outcome: | The proposed dataset covers the time periods related to pre- and post-war events, oil and gas discovery in Norway, and technological developments. |
EDEN: A Dataset for Event Detection in Norwegian News (2024.lrec-main)
Copied to clipboard
Samia Touileb, Jeanett Murstad, Petter Mæhlum, Lubos Steskal, Lilja Charlotte Storset, Huiling You, Lilja Øvrelid
| Challenge: | EDEN is the first dataset annotated with event information at the sentence level for the Norwegian language. |
| Approach: | They propose to annotate Norwegian news text and transcribed speech using ACE event schema. |
| Outcome: | The proposed dataset is the first annotated dataset for Norwegian, with a language-specific annotation process. |