| Challenge: | The first unidirectional parallel corpus spanish-croatia was built at the Faculty of Humanities and Social Sciences of the University of Zagreb. |
| Approach: | They describe the building of the first Spanish-Croatian unidirectional parallel corpus at the Faculty of Humanities and Social Sciences of the University of Zagreb. |
| Outcome: | The proposed corpus is a bilingual unidirectional (SpanishCroatian) parallel corpus . it contains 11 Spanish novels and their translations to Croatian done by six translators . |
Similar Papers
The EuroPat Corpus: A Parallel Corpus of European Patent Data (2022.lrec-1)
Copied to clipboard
| Challenge: | a new corpus of patent-specific parallel data is available for 6 official European languages paired with English: German, Spanish, French, Croatian, Norwegian, and Polish. |
| Approach: | They present a patent-specific corpus of parallel data for 6 official European languages paired with English: German, Spanish, French, Croatian, Norwegian, and Polish. |
| Outcome: | The filtered corpus ranges in size from 51 million sentences (Spanish-English) to 154k sentences (Croatian-English), with the unfiltered (raw) corpus being up to 2 times larger. |
Jojajovai: A Parallel Guarani-Spanish Corpus for MT Benchmarking (2022.lrec-1)
Copied to clipboard
Luis Chiruzzo, Santiago Góngora, Aldo Alvarez, Gustavo Giménez-Lugo, Marvin Agüero-Torales, Yliana Rodríguez
| Challenge: | a corpus of Guarani-Spanish text is presented that is aligned at sentence level . the long history of language contact between Guaran and Spanish in South America has resulted in many interesting language varieties . |
| Approach: | They propose to align Guarani-Spanish text at sentence level with 30,000 sentence pairs and a test set. |
| Outcome: | The proposed corpus contains about 30,000 sentence pairs and is structured as a collection of subsets from different sources, further split into training, development and test sets. |
A Large Parallel Corpus of Full-Text Scientific Articles (L18-1)
Copied to clipboard
| Challenge: | Scielo database contains articles from several research domains. |
| Approach: | They propose to build a parallel corpus from Scielo in three languages: English, Portuguese, and Spanish. |
| Outcome: | The proposed system outperforms other systems on scientific articles in English, Portuguese, and Spanish. |
DiHuTra: a Parallel Corpus to Analyse Differences between Human Translations (2022.lrec-1)
Copied to clipboard
| Challenge: | a new corpus of human translations contains both professional and student translations of news and reviews texts. |
| Approach: | They propose to use the data to compare human and professional translations of news and reviews in a new corpus which contains both professional and student translations. |
| Outcome: | The proposed corpus contains professional and student translations of news and reviews and a subcorpus containing reviews into Finnish. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
A Multilingual Parallel Corpora Collection Effort for Indian Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency. |
| Approach: | They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages. |
| Outcome: | The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health). |
A Multilingual Parallel Corpus for Aromanian (2024.lrec-main)
Copied to clipboard
| Challenge: | Aromanian is an endangered 1 language that currently lacks corpora and electronic resources that can potentially contribute to the preservation of its cultural heritage. |
| Approach: | They propose to create a corpus of Aromanian and equivalent sentence-aligned translations into Romanian, English, and French using orthographic standards. |
| Outcome: | The authors report that the first high-quality corpus of Aromanian is available in the Balkans and is available for download in Romanian, English, and French. |
C-XNLI: Croatian Extension of XNLI Dataset (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing parallel datasets limit multilingual evaluations due to the nature of linguistic annotation, which is tedious, subjective, and costly. |
| Approach: | They extend the Cross-lingual Natural Language Inference corpus with Croatian and use Facebook's 1.2B parameter m2m_100 model to analyze the train set and compare its quality with the existing machine-translated German set. |
| Outcome: | The proposed model is consistent with other XNLI dubs and is compared with the existing machine-translated German train set. |
FooTweets: A Bilingual Parallel Corpus of World Cup Tweets (L18-1)
Copied to clipboard
| Challenge: | a new study analyzes the nature of twitter data and compares it with other social networking websites. |
| Approach: | They develop a parallel corpus of tweets for an English-German pair that can be translated into German using a machine translation tool. |
| Outcome: | The proposed method can be used to translate tweets from English to German using a parallel corpus of 4, 000 tweets. |
VietMix: A Naturally-Occurring Parallel Corpus and Augmentation Framework for Vietnamese-English Code-Mixed Machine Translation (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing approaches to machine translation (MT) systems degrade when faced with code-mixed text. |
| Approach: | They propose a system that can augment Vietnamese-English code-mixed text with iterative fine-tuning and targeted filtering. |
| Outcome: | The proposed framework outperforms strong back-translation baselines and improves zero-shot models by up to +11.9 points. |