Challenge: a crowdsourcing project aimed at language learners has created a paraphrase corpus for 73 languages . the corpus contains 1.9 million sentences, with 200 - 250 000 sentences per language .
Approach: They propose to use a Tatoeba-based dataset to create a paraphrase corpus for 73 languages.
Outcome: The proposed dataset contains 1.9 million sentences and 200 - 250 000 sentences per language.

Similar Papers

Open Subtitles Paraphrase Corpus for Six Languages (L18-1)

Copied to clipboard

Challenge: Opusparcus is a new corpus of paraphrases for six European languages . it is based on movie and TV subtitles, which are colloquial and informal .
Approach: They propose to use opensubtitles2016 paraphrase corpus for six European languages . they extract paraphrases from movie and TV subtitles from the corpus .
Outcome: The new corpus is available in German, English, Finnish, French, Russian, and Swedish . it is extracted from the OpenSubtitles2016 corpus, which contains subtitles from movies and TV shows .
A Large Resource of Patterns for Verbal Paraphrases (L18-1)

Copied to clipboard

Challenge: Xu et al., 2015: paraphrases play an important role in natural language understanding . he says it is difficult to propose a paraphrasing relation for natural language processing systems .
Approach: They propose a resource of such paraphrases that can be used to identify hidden paraphrase pairs . they propose to use the resource to identify paraphrase relationships between two words .
Outcome: The proposed resource contains tens of thousands of such pairs and is available for academic purposes.
Multilingual Whispers: Generating Paraphrases with Translation (D19-55)

Copied to clipboard

Challenge: Humans naturally paraphrase, but they can generate approximately the same meaning with a different surface realization.
Approach: They compare translation-based paraphrase gathering using human, automatic, or hybrid techniques to monolingual paraphrasing by experts and non-experts.
Outcome: The proposed methods outperform human translation systems in a variety of translation tasks.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
ParaNMT-50M: Pushing the Limits of Paraphrastic Sentence Embeddings with Millions of Machine Translations (P18-1)

Copied to clipboard

Challenge: Using neural machine translation, we generate more than 50 million sentential paraphrase pairs from a large parallel corpus.
Approach: They use a dataset of more than 50 million English-English sentential paraphrase pairs to generate them automatically using neural machine translation.
Outcome: The proposed dataset outperforms all supervised systems on every SemEval semantic textual similarity competition and shows how it can be used for paraphrase generation.
ParaSCI: A Large Scientific Paraphrase Dataset for Longer Paraphrase Generation (2021.eacl-main)

Copied to clipboard

Challenge: Existing paraphrase datasets are mainly from news, novels, or social media platforms.
Approach: They propose to build a large-scale paraphrase dataset using intra-paper and inter-paper methods . they use PDBERT as a general paraphrase discovering method to take advantage of paraphrased sentences .
Outcome: The proposed dataset includes 33,981 paraphrase pairs from ACL and 316,063 pairs from arXiv . the major advantages of paraphrases lie in the prominent length and textual diversity .
Taxi1500: A Dataset for Multilingual Text Classification in 1500 Languages (2025.naacl-short)

Copied to clipboard

Challenge: a large-scale text classification dataset encompassing 1504 languages is needed to address this gap . low-resource languages are often overlooked due to the scarcity of evaluation datasets.
Approach: They propose to use translations of the Bible to construct a large-scale text classification dataset that covers 1504 languages and annotate them using crowdsourcing.
Outcome: The proposed dataset covers 1504 languages and is available to the public.
PAWS-X: A Cross-lingual Adversarial Dataset for Paraphrase Identification (D19-1)

Copied to clipboard

Challenge: Existing work on adversarial data generation focuses on English . Existing multilingual datasets show effectiveness of deep, multilingual pre-training .
Approach: They propose a dataset of 23,659 human translated PAWS evaluation pairs in six languages . they show the effectiveness of deep, multilingual pre-training while leaving considerable headroom .
Outcome: The proposed model shows that multilingual training and evaluation regimes are more accurate than previous models.
Using Paraphrases to Study Properties of Contextual Embeddings (2022.naacl-main)

Copied to clipboard

Challenge: Previously, paraphrases have been used to probe whether compositionality is accurately captured by BERT, but we believe they can be used to explore many other questions.
Approach: They propose to use paraphrases as a unique source of data to analyze contextualized embeddings, with a particular focus on BERT.
Outcome: The proposed analysis of paraphrases and paraphrase representations using the Paraphrase Database shows that BERT handles polysemous words, but different representations in many cases.
Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages (2020.lrec-1)

Copied to clipboard

Challenge: Using morphosyntactic tools, we train and distribute tools for approximately one thousand languages.
Approach: They train and distribute morphosyntactic tools for approximately one thousand languages.
Outcome: The results show that the tools generalize well across rare and common forms alike.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations