PAWS-X: A Cross-lingual Adversarial Dataset for Paraphrase Identification (D19-1)
Copied to clipboard
| Challenge: | Existing work on adversarial data generation focuses on English . Existing multilingual datasets show effectiveness of deep, multilingual pre-training . |
| Approach: | They propose a dataset of 23,659 human translated PAWS evaluation pairs in six languages . they show the effectiveness of deep, multilingual pre-training while leaving considerable headroom . |
| Outcome: | The proposed model shows that multilingual training and evaluation regimes are more accurate than previous models. |
Similar Papers
PAWS: Paraphrase Adversaries from Word Scrambling (N19-1)
Copied to clipboard
| Challenge: | Existing paraphrase identification datasets lack sentence pairs with high word overlap without being paraphrases. |
| Approach: | They propose a workflow for generating pairs of sentences with high word overlap . they use controlled word swapping and back translation followed by fluency and paraphrase judgments . |
| Outcome: | The proposed dataset has 108,463 well-formed paraphrase and non-paraphrase pairs with high lexical overlap. |
RuPAWS: A Russian Adversarial Dataset for Paraphrase Identification (2022.lrec-1)
Copied to clipboard
Nikita Martynov, Irina Krotova, Varvara Logacheva, Alexander Panchenko, Olga Kozlova, Nikita Semenov
| Challenge: | Existing datasets for paraphrase identification lack challenging sentence pairs with high word overlap. |
| Approach: | They propose to use a dataset for Russian paraphrase detection that includes examples from PAWS translated to the Russian language and manually annotated by native speakers. |
| Outcome: | The proposed model performs well on both datasets while maintaining accuracy on the ParaPhraser benchmark. |
Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Cross-Lingual Semantic Discrimination (CLSD) is a lightweight evaluation task that requires only parallel sentences and a Large Language Model (LLM) to generate adversarial distractors. |
| Approach: | They propose a lightweight task that requires only parallel sentences and a Large Language Model (LLM) to generate adversarial distractors. |
| Outcome: | The proposed task requires only parallel sentences and a Large Language Model (LLM) to generate adversarial distractors. |
ParaAMR: A Large-Scale Syntactically Diverse Paraphrase Dataset by AMR Back-Translation (2023.acl-long)
Copied to clipboard
| Challenge: | Paraphrase generation is a long-standing task in natural language processing (NLP). |
| Approach: | They propose to generate large-scale syntactically diverse paraphrase datasets by abstract meaning representation back-translation. |
| Outcome: | The proposed dataset is syntactically more diverse than existing datasets while maintaining good semantic similarity. |
Multilingual Whispers: Generating Paraphrases with Translation (D19-55)
Copied to clipboard
| Challenge: | Humans naturally paraphrase, but they can generate approximately the same meaning with a different surface realization. |
| Approach: | They compare translation-based paraphrase gathering using human, automatic, or hybrid techniques to monolingual paraphrasing by experts and non-experts. |
| Outcome: | The proposed methods outperform human translation systems in a variety of translation tasks. |
Bridging the Gap between Native Text and Translated Text through Adversarial Learning: A Case Study on Cross-Lingual Event Extraction (2023.findings-eacl)
Copied to clipboard
| Challenge: | Recent research in cross-lingual learning has found that combining large-scale pretrained multilingual language models with machine translation can yield good performance. |
| Approach: | They propose a model architecture that jointly encodes a source language input sentence with its translation to the target language during training and takes a target language sentence with it as input during evaluation. |
| Outcome: | The proposed model architecture can integrate machine translation to improve event extraction while adding machine-translated data yields unstable performance due to representational gap. |
Paraphrastic Representations at Scale (2022.emnlp-demos)
Copied to clipboard
| Challenge: | a new system allows users to train their own state-of-the-art paraphrastic sentence representations in a variety of languages. |
| Approach: | They propose a system that allows users to train their own paraphrastic sentence representations in a variety of languages. |
| Outcome: | The proposed models outperform previous models on monolingual and cross-lingual tasks and can be used on CPUs with little difference in inference speed. |
Improving Paraphrase Detection with the Adversarial Paraphrasing Task (2021.acl-long)
Copied to clipboard
| Challenge: | a new adversarial method of paraphrase identification is being used to identify paraphrases based on word overlap and syntax . authors propose a dataset that generates semantically equivalent but lexically and syntactically disparate paraphrase pairs . |
| Approach: | They propose an adversarial method for paraphrase identification that uses word overlap and syntax to identify paraphrases. |
| Outcome: | The proposed method improves paraphrase detection accuracy and speed of generation of datasets. |
Towards more equitable question answering systems: How much more data do you need? (2021.acl-short)
Copied to clipboard
| Challenge: | Question answering datasets in English are relatively new, but lack of linguistic diversity in the field is a challenge. |
| Approach: | They propose to use translation and cross-lingual transfer to produce QA systems in multiple languages to improve their performance. |
| Outcome: | The proposed approaches take advantage of existing resources to produce QA systems in multiple languages. |
Multilingual Offensive Language Identification with Cross-lingual Embeddings (2020.emnlp-main)
Copied to clipboard
| Challenge: | Several studies investigating methods to detect offensive content in social media use English data. |
| Approach: | They apply cross-lingual contextual embeddings and transfer learning to make predictions in languages with less resources. |
| Outcome: | The proposed method compares favorably to the best systems submitted to recent shared tasks on Bengali, Hindi, and Spanish. |