Papers by Mārcis Pinnis

6 papers
Tilde MT Platform for Developing Client Specific MT Solutions (L18-1)

Copied to clipboard

Challenge: a growing demand for translations and multilingual content is surpassing the supply of professional translation services.
Approach: They present a custom machine translation platform called Tilde MT that provides linguistic data storage, data cleaning and normalisation, statistical and neural machine translation system training and hosting functionality.
Outcome: The proposed platform provides linguistic data storage, data cleaning and normalisation, statistical and neural machine translation system training and hosting functionality, and wide integration capabilities.
Training and Adapting Multilingual NMT for Less-resourced and Morphologically Rich Languages (L18-1)

Copied to clipboard

Challenge: Using multilingual and multi-way neural machine translation approaches is a major advantage . training NMT systems for individual language pairs takes significantly more time than training of SMT systems .
Approach: They propose to employ multilingual and multi-way neural machine translation approaches for morphologically rich languages such as Estonian and Russian.
Outcome: The proposed approach improves translation quality by +3.27 BLEU points over baseline models.
Unsupervised Machine Translation in Real-World Scenarios (2022.lrec-1)

Copied to clipboard

Challenge: a recent study has shown that unsupervised methods rely on monolingual corpora to build MT systems.
Approach: They present the results of the MT4All CEF project using monolingual corpora . they propose to generate bilingual dictionaries and translation models from monolingual data .
Outcome: The proposed method generates bilingual dictionaries and translation models from monolingual corpora . results show that it is comparable to general domain supervised translation .
Code-Mixed Text Augmentation for Latvian ASR (2024.lrec-main)

Copied to clipboard

Challenge: a new study attempts to tackle code-mixed speech recognition by improving the language model of a hybrid system.
Approach: They propose an inflected transliteration and phonetic transcription model for code-mixed Latvian sentences . they leverage a large human-translated English-Latvian parallel text corpus to generate synthetic Latvian phrases .
Outcome: The proposed system improves on a human-translated English-Latvian parallel text corpus . the results show that the proposed system can generate code-mixed Latvian sentences .
Open Terminology Management and Sharing Toolkit for Federation of Terminology Databases (2022.lrec-1)

Copied to clipboard

Challenge: Terminology is also needed in AI applications such as machine translation, speech recognition, information extraction, and other natural language processing tools.
Approach: They propose a terminology management solution that facilitates standards-based sharing and management of terminology resources by providing the EuroTermBank Toolkit.
Outcome: The EuroTermBank Toolkit facilitates standards-based sharing and management of terminology resources by participating in federated databases.
Facilitating Terminology Translation with Target Lemma Annotations (2021.eacl-main)

Copied to clipboard

Challenge: Recent work on terminology integration assumes that the correct morphological forms are apriori known.
Approach: They propose to train machine translation systems using a source-side data augmentation method that annotates randomly selected source language words with their target language lemmas.
Outcome: The proposed method improves terminology translation accuracy in Latvian and Baltic languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations