Challenge: Existing parallel corpora for patents and scientific texts are not available due to the need for correct alignment and human curation.
Approach: They develop a parallel corpus from the open access Google Patents dataset . they use Hunalign algorithm to align sentences and tokens using the largest 22 languages .
Outcome: The proposed corpus is available in TSV format and with a SQLite database, with complementary information regarding patent metadata.

Similar Papers

JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus (2024.lrec-main)

Copied to clipboard

Challenge: a recent study has demonstrated that patent translation accuracy improves as the amount of training data or the number of model parameters increases.
Approach: They construct a bilingual corpus of Japanese-English patent application data from 2000 to 2021 . they extracted 1.4M Japanese- English document pairs and extracted 350M sentence pairs .
Outcome: The proposed method improves translation accuracy by 20 bleu points . it is the first publicly available large-scale Japanese-English patent corpus .
The EuroPat Corpus: A Parallel Corpus of European Patent Data (2022.lrec-1)

Copied to clipboard

Challenge: a new corpus of patent-specific parallel data is available for 6 official European languages paired with English: German, Spanish, French, Croatian, Norwegian, and Polish.
Approach: They present a patent-specific corpus of parallel data for 6 official European languages paired with English: German, Spanish, French, Croatian, Norwegian, and Polish.
Outcome: The filtered corpus ranges in size from 51 million sentences (Spanish-English) to 154k sentences (Croatian-English), with the unfiltered (raw) corpus being up to 2 times larger.
A Large Parallel Corpus of Full-Text Scientific Articles (L18-1)

Copied to clipboard

Challenge: Scielo database contains articles from several research domains.
Approach: They propose to build a parallel corpus from Scielo in three languages: English, Portuguese, and Spanish.
Outcome: The proposed system outperforms other systems on scientific articles in English, Portuguese, and Spanish.
Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 Indic Languages (2022.tacl-1)

Copied to clipboard

Challenge: We present Samanantar, the largest publicly available parallel corpora collection for Indic languages . based on existing corporative, there has been limited benefit for resource-poor languages despite the lack of parallel corporals and monolingual corporata.
Approach: They compile 12.4 million sentence pairs from existing corpora and mine 37.4 million from the Web.
Outcome: The proposed model outperforms existing models and benchmarks on public datasets.
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)

Copied to clipboard

Challenge: BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese .
Approach: They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages .
Outcome: The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese .
WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from Wikipedia (2021.eacl-main)

Copied to clipboard

Challenge: a new approach to extract parallel sentences from Wikipedia articles is proposed . the approach is based on multilingual sentence embeddings, but does not limit it to English .
Approach: They propose to automatically extract parallel sentences from Wikipedia articles in 96 languages . they train neural MT baseline systems on the mined data and evaluate them on the TED corpus .
Outcome: The proposed approach extracts parallel sentences from Wikipedia articles in 96 languages . the extracted sentences achieve strong BLEU scores for many language pairs .
Jojajovai: A Parallel Guarani-Spanish Corpus for MT Benchmarking (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of Guarani-Spanish text is presented that is aligned at sentence level . the long history of language contact between Guaran and Spanish in South America has resulted in many interesting language varieties .
Approach: They propose to align Guarani-Spanish text at sentence level with 30,000 sentence pairs and a test set.
Outcome: The proposed corpus contains about 30,000 sentence pairs and is structured as a collection of subsets from different sources, further split into training, development and test sets.
SciPar: A Collection of Parallel Corpora from Scientific Abstracts (2022.lrec-1)

Copied to clipboard

Challenge: SciPar is a collection of parallel corpora created from openly available metadata of bachelor theses, master theses and doctoral dissertations hosted in institutional repositories, digital libraries and national archives.
Approach: They propose to harvest and process openly available metadata from repositories to extract bilingual titles and abstracts from scientific publications.
Outcome: The proposed corpora could be useful for cross-lingual plagiarism detection or adapting Machine Translation systems for translation of scientific texts and academic writing in general.
SansTib, a Sanskrit - Tibetan Parallel Corpus and Bilingual Sentence Embedding Model (2022.lrec-1)

Copied to clipboard

Challenge: a large digital monolingual corpus of Sanskrit and Tibetan Buddhist literature has become available.
Approach: They propose to develop a Sanskrit - Classical Tibetan parallel corpus automatically aligned on sentence-level and a bilingual sentence embedding model.
Outcome: The proposed model improves the existing Sanskrit - Classical Tibetan parallel corpus and its bilingual sentence embedding model.
WikiAtomicEdits: A Multilingual Corpus of Wikipedia Edits for Modeling Language and Discourse (D18-1)

Copied to clipboard

Challenge: a corpus of 43 million atomic edits is available for Wikipedia edit history . edits are instances in which a human editor has inserted a single contiguous phrase into, or deleted a contigous phrase from, an existing sentence.
Approach: They use Wikipedia edit history to mine atomic edits across 8 languages . they find edits contain instances in which a human editor has inserted a single phrase into, or deleted a contiguous phrase from, an existing sentence.
Outcome: The data show that edits differ from the language observed in standard corpora and that models trained on edits encode different aspects of semantics and discourse than models trained in raw text.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations