Papers by Mārcis Pinnis
Tilde MT Platform for Developing Client Specific MT Solutions (L18-1)
Copied to clipboard
| Challenge: | a growing demand for translations and multilingual content is surpassing the supply of professional translation services. |
| Approach: | They present a custom machine translation platform called Tilde MT that provides linguistic data storage, data cleaning and normalisation, statistical and neural machine translation system training and hosting functionality. |
| Outcome: | The proposed platform provides linguistic data storage, data cleaning and normalisation, statistical and neural machine translation system training and hosting functionality, and wide integration capabilities. |
Training and Adapting Multilingual NMT for Less-resourced and Morphologically Rich Languages (L18-1)
Copied to clipboard
| Challenge: | Using multilingual and multi-way neural machine translation approaches is a major advantage . training NMT systems for individual language pairs takes significantly more time than training of SMT systems . |
| Approach: | They propose to employ multilingual and multi-way neural machine translation approaches for morphologically rich languages such as Estonian and Russian. |
| Outcome: | The proposed approach improves translation quality by +3.27 BLEU points over baseline models. |
Unsupervised Machine Translation in Real-World Scenarios (2022.lrec-1)
Copied to clipboard
Ona de Gibert Bonet, Iakes Goenaga, Jordi Armengol-Estapé, Olatz Perez-de-Viñaspre, Carla Parra Escartín, Marina Sanchez, Mārcis Pinnis, Gorka Labaka, Maite Melero
| Challenge: | a recent study has shown that unsupervised methods rely on monolingual corpora to build MT systems. |
| Approach: | They present the results of the MT4All CEF project using monolingual corpora . they propose to generate bilingual dictionaries and translation models from monolingual data . |
| Outcome: | The proposed method generates bilingual dictionaries and translation models from monolingual corpora . results show that it is comparable to general domain supervised translation . |
Code-Mixed Text Augmentation for Latvian ASR (2024.lrec-main)
Copied to clipboard
| Challenge: | a new study attempts to tackle code-mixed speech recognition by improving the language model of a hybrid system. |
| Approach: | They propose an inflected transliteration and phonetic transcription model for code-mixed Latvian sentences . they leverage a large human-translated English-Latvian parallel text corpus to generate synthetic Latvian phrases . |
| Outcome: | The proposed system improves on a human-translated English-Latvian parallel text corpus . the results show that the proposed system can generate code-mixed Latvian sentences . |
Open Terminology Management and Sharing Toolkit for Federation of Terminology Databases (2022.lrec-1)
Copied to clipboard
| Challenge: | Terminology is also needed in AI applications such as machine translation, speech recognition, information extraction, and other natural language processing tools. |
| Approach: | They propose a terminology management solution that facilitates standards-based sharing and management of terminology resources by providing the EuroTermBank Toolkit. |
| Outcome: | The EuroTermBank Toolkit facilitates standards-based sharing and management of terminology resources by participating in federated databases. |
Facilitating Terminology Translation with Target Lemma Annotations (2021.eacl-main)
Copied to clipboard
| Challenge: | Recent work on terminology integration assumes that the correct morphological forms are apriori known. |
| Approach: | They propose to train machine translation systems using a source-side data augmentation method that annotates randomly selected source language words with their target language lemmas. |
| Outcome: | The proposed method improves terminology translation accuracy in Latvian and Baltic languages. |