Papers by Veronika Laippala

7 papers
Beyond the English Web: Zero-Shot Cross-Lingual and Lightweight Monolingual Classification of Registers (2021.eacl-srw)

Copied to clipboard

Challenge: Existing studies on register classification for web documents have limited results due to skewed datasets and low performance.
Approach: They propose two new register-annotated corpora for French and Swedish . they show that deep pre-trained language models perform strongly in these languages .
Outcome: The proposed models outperform existing models in English and Finnish and can match or surpass existing models.
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT) (2025.acl-long)

Copied to clipboard

Challenge: a large number of textual data is needed to train state-of-the-art large language models.
Approach: They propose a collection of monolingual and parallel corpora from the Internet Archive . they document the entire data pipeline and release the code to reproduce it .
Outcome: The proposed collection of monolingual and parallel corpora is based on the HPLT v2 dataset . it includes 8T tokens covering 193 languages and 380M sentence pairs covering 51 languages .
Explaining Classes through Stable Word Attributions (2022.findings-acl)

Copied to clipboard

Challenge: Input saliency methods have become popular for explaining predictions of deep learning models, but there has been little work investigating methods for aggregating prediction-level explanations to the class level.
Approach: They propose a method to aggregate prediction-level explanations to the class level using XLM-R and Integrated Gradients input attribution methods.
Outcome: The proposed method extracts keyword lists of classes from text classification tasks and evaluates them on web register data.
Building Question-Answer Data Using Web Register Identification (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in web register (genre) identification have created a shortage of QA datasets for English and Finnish.
Approach: They propose a machine learning-based method for extracting QA pairs from web-scale data using XLM-R and a multilingual CORE web register corpus . they then develop a NER-style token classifier to identify the QA text spans within these documents.
Outcome: The proposed method is adaptable to any language given the availability of language models and extensive web data, but it is limited to English and Finnish.
A Broad-coverage Corpus for Finnish Named Entity Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Named entity recognition (NER) is a fundamental task in natural language processing (NLP).
Approach: They propose to annotate Finnish named entity names using a new corpus built on the Universal Dependencies corpus.
Outcome: The new annotation identifies over 10,000 mentions and maintains compatibility with a previously released single-domain corpus for Finnish NER.
FinGPT: Large Generative Models for a Small Language (2023.emnlp-main)

Copied to clipboard

Challenge: Neural language models excel in many tasks in NLP but are limited to smaller languages.
Approach: They propose two approaches to pretrain large language models for Finnish . they train seven monolingual models from scratch and use Finnish as pretraining data .
Outcome: The proposed model is based on a dataset of Finnish web crawls, news, social media and eBooks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations