Challenge: Existing work on adversarial data generation focuses on English . Existing multilingual datasets show effectiveness of deep, multilingual pre-training .
Approach: They propose a dataset of 23,659 human translated PAWS evaluation pairs in six languages . they show the effectiveness of deep, multilingual pre-training while leaving considerable headroom .
Outcome: The proposed model shows that multilingual training and evaluation regimes are more accurate than previous models.

Similar Papers

PAWS: Paraphrase Adversaries from Word Scrambling (N19-1)

Copied to clipboard

Challenge: Existing paraphrase identification datasets lack sentence pairs with high word overlap without being paraphrases.
Approach: They propose a workflow for generating pairs of sentences with high word overlap . they use controlled word swapping and back translation followed by fluency and paraphrase judgments .
Outcome: The proposed dataset has 108,463 well-formed paraphrase and non-paraphrase pairs with high lexical overlap.
RuPAWS: A Russian Adversarial Dataset for Paraphrase Identification (2022.lrec-1)

Copied to clipboard

Challenge: Existing datasets for paraphrase identification lack challenging sentence pairs with high word overlap.
Approach: They propose to use a dataset for Russian paraphrase detection that includes examples from PAWS translated to the Russian language and manually annotated by native speakers.
Outcome: The proposed model performs well on both datasets while maintaining accuracy on the ParaPhraser benchmark.
Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples (2025.findings-emnlp)

Copied to clipboard

Challenge: Cross-Lingual Semantic Discrimination (CLSD) is a lightweight evaluation task that requires only parallel sentences and a Large Language Model (LLM) to generate adversarial distractors.
Approach: They propose a lightweight task that requires only parallel sentences and a Large Language Model (LLM) to generate adversarial distractors.
Outcome: The proposed task requires only parallel sentences and a Large Language Model (LLM) to generate adversarial distractors.
ParaAMR: A Large-Scale Syntactically Diverse Paraphrase Dataset by AMR Back-Translation (2023.acl-long)

Copied to clipboard

Challenge: Paraphrase generation is a long-standing task in natural language processing (NLP).
Approach: They propose to generate large-scale syntactically diverse paraphrase datasets by abstract meaning representation back-translation.
Outcome: The proposed dataset is syntactically more diverse than existing datasets while maintaining good semantic similarity.
Multilingual Whispers: Generating Paraphrases with Translation (D19-55)

Copied to clipboard

Challenge: Humans naturally paraphrase, but they can generate approximately the same meaning with a different surface realization.
Approach: They compare translation-based paraphrase gathering using human, automatic, or hybrid techniques to monolingual paraphrasing by experts and non-experts.
Outcome: The proposed methods outperform human translation systems in a variety of translation tasks.
Bridging the Gap between Native Text and Translated Text through Adversarial Learning: A Case Study on Cross-Lingual Event Extraction (2023.findings-eacl)

Copied to clipboard

Challenge: Recent research in cross-lingual learning has found that combining large-scale pretrained multilingual language models with machine translation can yield good performance.
Approach: They propose a model architecture that jointly encodes a source language input sentence with its translation to the target language during training and takes a target language sentence with it as input during evaluation.
Outcome: The proposed model architecture can integrate machine translation to improve event extraction while adding machine-translated data yields unstable performance due to representational gap.
Paraphrastic Representations at Scale (2022.emnlp-demos)

Copied to clipboard

Challenge: a new system allows users to train their own state-of-the-art paraphrastic sentence representations in a variety of languages.
Approach: They propose a system that allows users to train their own paraphrastic sentence representations in a variety of languages.
Outcome: The proposed models outperform previous models on monolingual and cross-lingual tasks and can be used on CPUs with little difference in inference speed.
Improving Paraphrase Detection with the Adversarial Paraphrasing Task (2021.acl-long)

Copied to clipboard

Challenge: a new adversarial method of paraphrase identification is being used to identify paraphrases based on word overlap and syntax . authors propose a dataset that generates semantically equivalent but lexically and syntactically disparate paraphrase pairs .
Approach: They propose an adversarial method for paraphrase identification that uses word overlap and syntax to identify paraphrases.
Outcome: The proposed method improves paraphrase detection accuracy and speed of generation of datasets.
Towards more equitable question answering systems: How much more data do you need? (2021.acl-short)

Copied to clipboard

Challenge: Question answering datasets in English are relatively new, but lack of linguistic diversity in the field is a challenge.
Approach: They propose to use translation and cross-lingual transfer to produce QA systems in multiple languages to improve their performance.
Outcome: The proposed approaches take advantage of existing resources to produce QA systems in multiple languages.
Multilingual Offensive Language Identification with Cross-lingual Embeddings (2020.emnlp-main)

Copied to clipboard

Challenge: Several studies investigating methods to detect offensive content in social media use English data.
Approach: They apply cross-lingual contextual embeddings and transfer learning to make predictions in languages with less resources.
Outcome: The proposed method compares favorably to the best systems submitted to recent shared tasks on Bengali, Hindi, and Spanish.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations