Papers by Miika Oinonen

2 papers
Beyond the English Web: Zero-Shot Cross-Lingual and Lightweight Monolingual Classification of Registers (2021.eacl-srw)

Copied to clipboard

Challenge: Existing studies on register classification for web documents have limited results due to skewed datasets and low performance.
Approach: They propose two new register-annotated corpora for French and Swedish . they show that deep pre-trained language models perform strongly in these languages .
Outcome: The proposed models outperform existing models in English and Finnish and can match or surpass existing models.
A Broad-coverage Corpus for Finnish Named Entity Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Named entity recognition (NER) is a fundamental task in natural language processing (NLP).
Approach: They propose to annotate Finnish named entity names using a new corpus built on the Universal Dependencies corpus.
Outcome: The new annotation identifies over 10,000 mentions and maintains compatibility with a previously released single-domain corpus for Finnish NER.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations