Papers by Jenna Kanerva

6 papers
Parse Me if You Can: Artificial Treebanks for Parsing Experiments on Elliptical Constructions (L18-1)

Copied to clipboard

Challenge: ellipsis is a phenomenon present in many natural languages, but it complicates syntactic parsing of the content that is not omitted.
Approach: They analyze outputs of state-of-the-art parsers to learn about parsing accuracy and typical errors from the perspective of elliptical constructions.
Outcome: The proposed treebank is a semi-artificially constructed treebank of ellipsis.
GEMv2: Multilingual NLG Benchmarking in a Single Line of Code (2022.emnlp-demos)

Copied to clipboard

Challenge: Evaluations in machine learning rarely use the latest metrics, datasets, or human evaluation in favor of remaining compatible with prior work.
Approach: They propose to use the Generation, Evaluation, and Metrics Benchmark to integrate new evaluation methods into existing evaluations.
Outcome: The proposed evaluation infrastructure bridges the gap between the advantages of leaderboards and in-depth and evolving evaluations by allowing model developers to benefit from each other's work.
The FISKMÖ Project: Resources and Tools for Finnish-Swedish Machine Translation and Cross-Linguistic Research (2020.lrec-1)

Copied to clipboard

Challenge: Finnish and Swedish are the two official languages of Finland.
Approach: They propose to compile a massive corpus of translated material between Finnish and Swedish . they also aim to develop open and freely accessible translation services for those two languages .
Outcome: The project aims to develop open and freely accessible translation services for Finnish and Swedish.
Out-of-Domain Evaluation of Finnish Dependency Parsing (2022.lrec-1)

Copied to clipboard

Challenge: prevailing practice in academia evaluates model performance on in-domain evaluation data . however, in many real world applications data on which model is applied may differ from training data - a problem that is not addressed by current literature.
Approach: They propose to use Finnish-OOD out-of-domain treebank for out- of-domain evaluation . they propose to include sections more challenging for the general parser .
Outcome: The proposed treebank includes five distinct data sources and a total of 19,382 syntactic words in 2,122 sentences.
FinGPT: Large Generative Models for a Small Language (2023.emnlp-main)

Copied to clipboard

Challenge: Neural language models excel in many tasks in NLP but are limited to smaller languages.
Approach: They propose two approaches to pretrain large language models for Finnish . they train seven monolingual models from scratch and use Finnish as pretraining data .
Outcome: The proposed model is based on a dataset of Finnish web crawls, news, social media and eBooks.
Neural Dependency Parsing of Biomedical Text: TurkuNLP entry in the CRAFT Structural Annotation Task (D19-57)

Copied to clipboard

Challenge: Syntactic analysis (parsing) is a fundamental task in natural language processing (NLP).
Approach: They propose to use the Turku neural parser to adapt it to the biomedical domain . they evaluated custom word embeddings, combination with other in-domain resources .
Outcome: The proposed approach achieved a labeled attachment score of 89.7%, the best among task participants.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations