| Challenge: | a crowdsourcing project aimed at language learners has created a paraphrase corpus for 73 languages . the corpus contains 1.9 million sentences, with 200 - 250 000 sentences per language . |
| Approach: | They propose to use a Tatoeba-based dataset to create a paraphrase corpus for 73 languages. |
| Outcome: | The proposed dataset contains 1.9 million sentences and 200 - 250 000 sentences per language. |
Similar Papers
Open Subtitles Paraphrase Corpus for Six Languages (L18-1)
Copied to clipboard
| Challenge: | Opusparcus is a new corpus of paraphrases for six European languages . it is based on movie and TV subtitles, which are colloquial and informal . |
| Approach: | They propose to use opensubtitles2016 paraphrase corpus for six European languages . they extract paraphrases from movie and TV subtitles from the corpus . |
| Outcome: | The new corpus is available in German, English, Finnish, French, Russian, and Swedish . it is extracted from the OpenSubtitles2016 corpus, which contains subtitles from movies and TV shows . |
A Large Resource of Patterns for Verbal Paraphrases (L18-1)
Copied to clipboard
| Challenge: | Xu et al., 2015: paraphrases play an important role in natural language understanding . he says it is difficult to propose a paraphrasing relation for natural language processing systems . |
| Approach: | They propose a resource of such paraphrases that can be used to identify hidden paraphrase pairs . they propose to use the resource to identify paraphrase relationships between two words . |
| Outcome: | The proposed resource contains tens of thousands of such pairs and is available for academic purposes. |
Multilingual Whispers: Generating Paraphrases with Translation (D19-55)
Copied to clipboard
| Challenge: | Humans naturally paraphrase, but they can generate approximately the same meaning with a different surface realization. |
| Approach: | They compare translation-based paraphrase gathering using human, automatic, or hybrid techniques to monolingual paraphrasing by experts and non-experts. |
| Outcome: | The proposed methods outperform human translation systems in a variety of translation tasks. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
ParaNMT-50M: Pushing the Limits of Paraphrastic Sentence Embeddings with Millions of Machine Translations (P18-1)
Copied to clipboard
| Challenge: | Using neural machine translation, we generate more than 50 million sentential paraphrase pairs from a large parallel corpus. |
| Approach: | They use a dataset of more than 50 million English-English sentential paraphrase pairs to generate them automatically using neural machine translation. |
| Outcome: | The proposed dataset outperforms all supervised systems on every SemEval semantic textual similarity competition and shows how it can be used for paraphrase generation. |
ParaSCI: A Large Scientific Paraphrase Dataset for Longer Paraphrase Generation (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing paraphrase datasets are mainly from news, novels, or social media platforms. |
| Approach: | They propose to build a large-scale paraphrase dataset using intra-paper and inter-paper methods . they use PDBERT as a general paraphrase discovering method to take advantage of paraphrased sentences . |
| Outcome: | The proposed dataset includes 33,981 paraphrase pairs from ACL and 316,063 pairs from arXiv . the major advantages of paraphrases lie in the prominent length and textual diversity . |
Taxi1500: A Dataset for Multilingual Text Classification in 1500 Languages (2025.naacl-short)
Copied to clipboard
| Challenge: | a large-scale text classification dataset encompassing 1504 languages is needed to address this gap . low-resource languages are often overlooked due to the scarcity of evaluation datasets. |
| Approach: | They propose to use translations of the Bible to construct a large-scale text classification dataset that covers 1504 languages and annotate them using crowdsourcing. |
| Outcome: | The proposed dataset covers 1504 languages and is available to the public. |
PAWS-X: A Cross-lingual Adversarial Dataset for Paraphrase Identification (D19-1)
Copied to clipboard
| Challenge: | Existing work on adversarial data generation focuses on English . Existing multilingual datasets show effectiveness of deep, multilingual pre-training . |
| Approach: | They propose a dataset of 23,659 human translated PAWS evaluation pairs in six languages . they show the effectiveness of deep, multilingual pre-training while leaving considerable headroom . |
| Outcome: | The proposed model shows that multilingual training and evaluation regimes are more accurate than previous models. |
Using Paraphrases to Study Properties of Contextual Embeddings (2022.naacl-main)
Copied to clipboard
| Challenge: | Previously, paraphrases have been used to probe whether compositionality is accurately captured by BERT, but we believe they can be used to explore many other questions. |
| Approach: | They propose to use paraphrases as a unique source of data to analyze contextualized embeddings, with a particular focus on BERT. |
| Outcome: | The proposed analysis of paraphrases and paraphrase representations using the Paraphrase Database shows that BERT handles polysemous words, but different representations in many cases. |
Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Using morphosyntactic tools, we train and distribute tools for approximately one thousand languages. |
| Approach: | They train and distribute morphosyntactic tools for approximately one thousand languages. |
| Outcome: | The results show that the tools generalize well across rare and common forms alike. |