| Challenge: | a new corpus of patent-specific parallel data is available for 6 official European languages paired with English: German, Spanish, French, Croatian, Norwegian, and Polish. |
| Approach: | They present a patent-specific corpus of parallel data for 6 official European languages paired with English: German, Spanish, French, Croatian, Norwegian, and Polish. |
| Outcome: | The filtered corpus ranges in size from 51 million sentences (Spanish-English) to 154k sentences (Croatian-English), with the unfiltered (raw) corpus being up to 2 times larger. |
Similar Papers
ParaPat: The Multi-Million Sentences Parallel Corpus of Patents Abstracts (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing parallel corpora for patents and scientific texts are not available due to the need for correct alignment and human curation. |
| Approach: | They develop a parallel corpus from the open access Google Patents dataset . they use Hunalign algorithm to align sentences and tokens using the largest 22 languages . |
| Outcome: | The proposed corpus is available in TSV format and with a SQLite database, with complementary information regarding patent metadata. |
JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | a recent study has demonstrated that patent translation accuracy improves as the amount of training data or the number of model parameters increases. |
| Approach: | They construct a bilingual corpus of Japanese-English patent application data from 2000 to 2021 . they extracted 1.4M Japanese- English document pairs and extracted 350M sentence pairs . |
| Outcome: | The proposed method improves translation accuracy by 20 bleu points . it is the first publicly available large-scale Japanese-English patent corpus . |
Building the Spanish-Croatian Parallel Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | The first unidirectional parallel corpus spanish-croatia was built at the Faculty of Humanities and Social Sciences of the University of Zagreb. |
| Approach: | They describe the building of the first Spanish-Croatian unidirectional parallel corpus at the Faculty of Humanities and Social Sciences of the University of Zagreb. |
| Outcome: | The proposed corpus is a bilingual unidirectional (SpanishCroatian) parallel corpus . it contains 11 Spanish novels and their translations to Croatian done by six translators . |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation (L18-1)
Copied to clipboard
| Challenge: | Recent work in spoken language translation (SLT) has attempted to build end-to-end speech-totext translation without using source language transcription during learning or decoding. |
| Approach: | They propose to augment an existing (monolingual) corpus: LibriSpeech. |
| Outcome: | The proposed corpus is derived from read audiobooks from the LibriVox project and has been carefully segmented and aligned. |
A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent studies have highlighted the potential of exploiting parallel corpora to enhance multilingual large language models. |
| Approach: | They investigate the impact of parallel corpora quality and quantity, training objectives, and model size on performance of multilingual large language models enhanced with parallel corporeal. |
| Outcome: | The proposed approach improves performance in bilingual and general-purpose tasks. |
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)
Copied to clipboard
| Challenge: | BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese . |
| Approach: | They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages . |
| Outcome: | The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese . |
Data Augmentation for Code Translation with Comparable Corpora and Multiple References (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for translating code between programming languages are limited by parallel training data. |
| Approach: | They propose a data augmentation technique that builds comparable corpora and augments existing parallel data with multiple reference translations. |
| Outcome: | The proposed techniques improve CodeT5 translation between Java, Python, and C++ by an average of 7.5% Computational Accuracy (CA@1) . |
WikiAtomicEdits: A Multilingual Corpus of Wikipedia Edits for Modeling Language and Discourse (D18-1)
Copied to clipboard
| Challenge: | a corpus of 43 million atomic edits is available for Wikipedia edit history . edits are instances in which a human editor has inserted a single contiguous phrase into, or deleted a contigous phrase from, an existing sentence. |
| Approach: | They use Wikipedia edit history to mine atomic edits across 8 languages . they find edits contain instances in which a human editor has inserted a single phrase into, or deleted a contiguous phrase from, an existing sentence. |
| Outcome: | The data show that edits differ from the language observed in standard corpora and that models trained on edits encode different aspects of semantics and discourse than models trained in raw text. |
The IIT Bombay English-Hindi Parallel Corpus (L18-1)
Copied to clipboard
| Challenge: | a corpus of 1.49 million parallel segments is available in the public domain . the corpus is the largest publicly available English-Hindi parallel corpus . |
| Approach: | They present the IIT Bombay English-Hindi Parallel Corpus . they present a compilation of public and private parallel corpora . |
| Outcome: | The corpus contains 1.49 million parallel segments, of which 694k were not previously available in the public domain. |