Papers by Miika Oinonen
Beyond the English Web: Zero-Shot Cross-Lingual and Lightweight Monolingual Classification of Registers (2021.eacl-srw)
Copied to clipboard
Liina Repo, Valtteri Skantsi, Samuel Rönnqvist, Saara Hellström, Miika Oinonen, Anna Salmela, Douglas Biber, Jesse Egbert, Sampo Pyysalo, Veronika Laippala
| Challenge: | Existing studies on register classification for web documents have limited results due to skewed datasets and low performance. |
| Approach: | They propose two new register-annotated corpora for French and Swedish . they show that deep pre-trained language models perform strongly in these languages . |
| Outcome: | The proposed models outperform existing models in English and Finnish and can match or surpass existing models. |
A Broad-coverage Corpus for Finnish Named Entity Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a fundamental task in natural language processing (NLP). |
| Approach: | They propose to annotate Finnish named entity names using a new corpus built on the Universal Dependencies corpus. |
| Outcome: | The new annotation identifies over 10,000 mentions and maintains compatibility with a previously released single-domain corpus for Finnish NER. |