The FISKMÖ Project: Resources and Tools for Finnish-Swedish Machine Translation and Cross-Linguistic Research (2020.lrec-1)
Copied to clipboard
Jörg Tiedemann, Tommi Nieminen, Mikko Aulamo, Jenna Kanerva, Akseli Leino, Filip Ginter, Niko Papula
| Challenge: | Finnish and Swedish are the two official languages of Finland. |
| Approach: | They propose to compile a massive corpus of translated material between Finnish and Swedish . they also aim to develop open and freely accessible translation services for those two languages . |
| Outcome: | The project aims to develop open and freely accessible translation services for Finnish and Swedish. |
Similar Papers
Language Model Priors and Data Augmentation Strategies for Low-resource Machine Translation: A Case Study Using Finnish to Northern Sámi (2024.findings-acl)
Copied to clipboard
| Challenge: | a new study examines the use of monolingual data for improving low-resource machine translation. |
| Approach: | They investigate ways of using monolingual data for improving low-resource machine translation. |
| Outcome: | The proposed model can perform better on the target-side data without augmentation of parallel data. |
Machine Translation for Low-Resource Languages through Monolingual Data and LLM: A Case Study of English-to-Basque (2026.eacl-srw)
Copied to clipboard
| Challenge: | Existing LLMs do not translate well from English to Basque, but they yield an acceptable performance in the reverse direction. |
| Approach: | They propose to use a Basque monolingual corpora to train an LLM-based MT system . they use 'sovereignty fine tuning' to generate parallel corporata, and then use preference optimization . |
| Outcome: | The proposed system improves translation quality in English-to-Basque direction while requiring limited data for low-resource languages. |
Unifying Cross-Lingual Transfer across Scenarios of Resource Scarcity (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to deal with resource scarcity have not been developed to deal effectively with the problem. |
| Approach: | They propose to use a set of tools to harness data from one or more high-resource "source" languages to compensate for a shortage of data in low-resourced "target" languages. |
| Outcome: | The proposed technique can be easily adapted to unseen languages, extending the range of the proposed technique and translation-based transfer more broadly. |
Many-to-English Machine Translation Tools, Data, and Pretrained Models (2021.acl-demo)
Copied to clipboard
| Challenge: | Commercial translation systems support only one hundred languages or fewer . commercial translation systems do not make these models available for transfer to low resource languages . |
| Approach: | They propose a multilingual neural machine translation model that can translate from 500 source languages to English. |
| Outcome: | The proposed model can translate from 500 source languages to English, or be used as a parent model for low-resource languages. |
Tilde MT Platform for Developing Client Specific MT Solutions (L18-1)
Copied to clipboard
| Challenge: | a growing demand for translations and multilingual content is surpassing the supply of professional translation services. |
| Approach: | They present a custom machine translation platform called Tilde MT that provides linguistic data storage, data cleaning and normalisation, statistical and neural machine translation system training and hosting functionality. |
| Outcome: | The proposed platform provides linguistic data storage, data cleaning and normalisation, statistical and neural machine translation system training and hosting functionality, and wide integration capabilities. |
compare-mt: A Tool for Holistic Comparison of Language Generation Systems (N19-4)
Copied to clipboard
| Challenge: | Unlike machine translation, natural language outputs are nuanced and there are no clear yes/no distinctions about whether they are correct or not. |
| Approach: | They describe compare-mt, a tool for holistic analysis and comparison of the results of systems for language generation tasks such as machine translation. |
| Outcome: | The compare-mt tool is an open-source pure-python package that has already proven useful to generate analyses that have been used in our papers. |
DiHuTra: a Parallel Corpus to Analyse Differences between Human Translations (2022.lrec-1)
Copied to clipboard
| Challenge: | a new corpus of human translations contains both professional and student translations of news and reviews texts. |
| Approach: | They propose to use the data to compare human and professional translations of news and reviews in a new corpus which contains both professional and student translations. |
| Outcome: | The proposed corpus contains professional and student translations of news and reviews and a subcorpus containing reviews into Finnish. |
FinGPT: Large Generative Models for a Small Language (2023.emnlp-main)
Copied to clipboard
Risto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen, Jenna Kanerva, Hanna-Mari Kupari, Filip Ginter, Veronika Laippala, Niklas Muennighoff, Aleksandra Piktus, Thomas Wang, Nouamane Tazi, Teven Scao, Thomas Wolf, Osma Suominen, Samuli Sairanen, Mikko Merioksa, Jyrki Heinonen, Aija Vahtola, Samuel Antao, Sampo Pyysalo
| Challenge: | Neural language models excel in many tasks in NLP but are limited to smaller languages. |
| Approach: | They propose two approaches to pretrain large language models for Finnish . they train seven monolingual models from scratch and use Finnish as pretraining data . |
| Outcome: | The proposed model is based on a dataset of Finnish web crawls, news, social media and eBooks. |
MIT-10M: A Large Scale Parallel Corpus of Multilingual Image Translation (2025.coling-main)
Copied to clipboard
| Challenge: | Existing datasets suffer from limitations in scale, diversity, and quality, hindering the development and evaluation of IT models. |
| Approach: | They propose a large-scale parallel corpus of multilingual image translation with over 10M image-text pairs derived from real-world data. |
| Outcome: | The proposed model performs better in tackling challenging and complex image translation tasks in the real world. |
Data Cartography for Low-Resource Neural Machine Translation (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to improve machine translation (MT) in low-resource settings are limited in the number of languages spoken in the world. |
| Approach: | They apply cartography techniques to characterize the contribution of training samples in two low-resource MT tasks (Swahili-English and Turkish-English) they argue that data augmentation strategies for low-Resource ML would benefit from model-in-the-loop strategies to maximize improvements. |
| Outcome: | The proposed methods show that training samples contribute to model training in low-resource MT tasks, albeit not uniformly throughout the training process. |