Building the Spanish-Croatian Parallel Corpus (2020.lrec-1)

Copied to clipboard

Challenge: The first unidirectional parallel corpus spanish-croatia was built at the Faculty of Humanities and Social Sciences of the University of Zagreb.
Approach: They describe the building of the first Spanish-Croatian unidirectional parallel corpus at the Faculty of Humanities and Social Sciences of the University of Zagreb.
Outcome: The proposed corpus is a bilingual unidirectional (SpanishCroatian) parallel corpus . it contains 11 Spanish novels and their translations to Croatian done by six translators .

Similar Papers

The EuroPat Corpus: A Parallel Corpus of European Patent Data (2022.lrec-1)

Copied to clipboard

Challenge: a new corpus of patent-specific parallel data is available for 6 official European languages paired with English: German, Spanish, French, Croatian, Norwegian, and Polish.
Approach: They present a patent-specific corpus of parallel data for 6 official European languages paired with English: German, Spanish, French, Croatian, Norwegian, and Polish.
Outcome: The filtered corpus ranges in size from 51 million sentences (Spanish-English) to 154k sentences (Croatian-English), with the unfiltered (raw) corpus being up to 2 times larger.
Jojajovai: A Parallel Guarani-Spanish Corpus for MT Benchmarking (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of Guarani-Spanish text is presented that is aligned at sentence level . the long history of language contact between Guaran and Spanish in South America has resulted in many interesting language varieties .
Approach: They propose to align Guarani-Spanish text at sentence level with 30,000 sentence pairs and a test set.
Outcome: The proposed corpus contains about 30,000 sentence pairs and is structured as a collection of subsets from different sources, further split into training, development and test sets.
A Large Parallel Corpus of Full-Text Scientific Articles (L18-1)

Copied to clipboard

Challenge: Scielo database contains articles from several research domains.
Approach: They propose to build a parallel corpus from Scielo in three languages: English, Portuguese, and Spanish.
Outcome: The proposed system outperforms other systems on scientific articles in English, Portuguese, and Spanish.
DiHuTra: a Parallel Corpus to Analyse Differences between Human Translations (2022.lrec-1)

Copied to clipboard

Challenge: a new corpus of human translations contains both professional and student translations of news and reviews texts.
Approach: They propose to use the data to compare human and professional translations of news and reviews in a new corpus which contains both professional and student translations.
Outcome: The proposed corpus contains professional and student translations of news and reviews and a subcorpus containing reviews into Finnish.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
A Multilingual Parallel Corpora Collection Effort for Indian Languages (2020.lrec-1)

Copied to clipboard

Challenge: Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency.
Approach: They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages.
Outcome: The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health).
A Multilingual Parallel Corpus for Aromanian (2024.lrec-main)

Copied to clipboard

Challenge: Aromanian is an endangered 1 language that currently lacks corpora and electronic resources that can potentially contribute to the preservation of its cultural heritage.
Approach: They propose to create a corpus of Aromanian and equivalent sentence-aligned translations into Romanian, English, and French using orthographic standards.
Outcome: The authors report that the first high-quality corpus of Aromanian is available in the Balkans and is available for download in Romanian, English, and French.
C-XNLI: Croatian Extension of XNLI Dataset (2023.findings-acl)

Copied to clipboard

Challenge: Existing parallel datasets limit multilingual evaluations due to the nature of linguistic annotation, which is tedious, subjective, and costly.
Approach: They extend the Cross-lingual Natural Language Inference corpus with Croatian and use Facebook's 1.2B parameter m2m_100 model to analyze the train set and compare its quality with the existing machine-translated German set.
Outcome: The proposed model is consistent with other XNLI dubs and is compared with the existing machine-translated German train set.
FooTweets: A Bilingual Parallel Corpus of World Cup Tweets (L18-1)

Copied to clipboard

Challenge: a new study analyzes the nature of twitter data and compares it with other social networking websites.
Approach: They develop a parallel corpus of tweets for an English-German pair that can be translated into German using a machine translation tool.
Outcome: The proposed method can be used to translate tweets from English to German using a parallel corpus of 4, 000 tweets.
VietMix: A Naturally-Occurring Parallel Corpus and Augmentation Framework for Vietnamese-English Code-Mixed Machine Translation (2026.eacl-long)

Copied to clipboard

Challenge: Existing approaches to machine translation (MT) systems degrade when faced with code-mixed text.
Approach: They propose a system that can augment Vietnamese-English code-mixed text with iterative fine-tuning and targeted filtering.
Outcome: The proposed framework outperforms strong back-translation baselines and improves zero-shot models by up to +11.9 points.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations