Challenge: Argument Reasoning Comprehension Task (ARCT) focuses on inferences, not just discovering warrants.
Approach: They propose to build an adversarial dataset on which all models achieve random accuracy.
Outcome: The proposed dataset provides a more robust assessment of argument comprehension and should be adopted as the standard in future work.

Similar Papers

Using Adversarial Attacks to Reveal the Statistical Bias in Machine Reading Comprehension Models (2021.acl-short)

Copied to clipboard

Challenge: Pre-trained language models have achieved human-level performance on many Machine Reading Comprehension (MRC) tasks, but it remains unclear whether these models truly understand language or answer questions by exploiting statistical biases in datasets.
Approach: They propose a method to attack MRC models by exposing statistical biases in a RACE dataset and propose an augmented training method that can greatly reduce models’ statistical bias.
Outcome: The proposed method can reduce models’ statistical biases from human-level performance to chance-level.
Listening Comprehension over Argumentative Content (D18-1)

Copied to clipboard

Challenge: In argumentation domain, people are exposed directly to audio (or the video), without access to a written version.
Approach: They present a task for machine listening comprehension in the argumentation domain and a dataset in English.
Outcome: The proposed task is based on 200 speeches arguing for or against 50 controversial topics and uses baseline methods to address it.
Adversarial NLI: A New Benchmark for Natural Language Understanding (2020.acl-main)

Copied to clipboard

Challenge: a new large-scale NLI benchmark dataset is presented to test models on a variety of popular NLIs.
Approach: They propose a large-scale NLI benchmark dataset that is iteratively compared with a human-and-model-in-the-loop procedure.
Outcome: The proposed method can be applied in a never-ending learning scenario, becoming a moving target for NLU, rather than a static benchmark that will quickly saturate.
Towards Interpreting BERT for Reading Comprehension Based QA (2020.emnlp-main)

Copied to clipboard

Challenge: Pretrained language models such as ELMO and XLNet have achieved state-of-the-art performance on various NLP tasks.
Approach: They propose to define a layer’s role or functionality using Integrated Gradients and perform preliminary analysis across all layers.
Outcome: The proposed model performs better than existing models on RCQA and ELMO, but it lacks the human-level performance needed to perform the task.
Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not Arguments (2025.acl-long)

Copied to clipboard

Challenge: Identifying arguments is a prerequisite for various tasks in automated discourse analysis.
Approach: They evaluate four BERT-like transformers on 17 English sentence-level datasets . they find that they tend to rely on lexical shortcuts tied to content words .
Outcome: The proposed models perform best on 17 English sentence-level datasets on common tasks, but their performance drops when applied to unseen datasets.
Giving BERT a Calculator: Finding Operations and Arguments with Reading Comprehension (D19-1)

Copied to clipboard

Challenge: End-to-end reading comprehension models have been successful at extracting text answers, but there are still problems with generalizing them to abstractive numerical reasoning.
Approach: They propose to augment a BERT-based reading comprehension model with a set of executable ‘programs’ which encompass simple arithmetic as well as extraction.
Outcome: The proposed model can perform 33% absolute improvement on the DROP dataset, with very few training examples.
Improving the Robustness of Deep Reading Comprehension Models by Leveraging Syntax Prior (D19-58)

Copied to clipboard

Challenge: Recent studies indicate that the current machine reading comprehension systems suffer from weak robustness against adversarial samples.
Approach: They propose to take sentence syntax as the leverage in the answer predicting process and exploit the syntactic elements of a question to improve the generalization and robustness of MRC models.
Outcome: The proposed method improves generalization and robustness against adversarial samples, with performance well-maintained.
Quoref: A Reading Comprehension Dataset with Questions Requiring Coreferential Reasoning (D19-1)

Copied to clipboard

Challenge: Existing reading comprehension benchmarks do not contain complex coreferential phenomena . obtaining questions focused on such phenomena is difficult because of lexical cues .
Approach: They propose to use a crowdsourced dataset to examine the ability of models to resolve coreference among entities in Wikipedia paragraphs.
Outcome: The proposed model performs significantly worse than humans on the reading comprehension benchmark . paragraphs and other longer texts typically make multiple references to the same entities .
Beat the AI: Investigating Adversarial Human Annotation for Reading Comprehension (2020.tacl-1)

Copied to clipboard

Challenge: Innovations in annotation methodologies have been a catalyst for Reading Comprehension (RC) datasets and models.
Approach: They propose to use a model-in-the-annotation-loop approach to train adversarial models in three different settings to explore reproducibility of the adversarial effect, transfer from data collected with varying model- in-the loop strengths, and generalization to data collected without a modeling model.
Outcome: The proposed approach achieves 39.9F1 on questions it cannot answer when trained on SQUAD, but lower than when trained using RoBERTa itself (41.0F1).
HellaSwag: Can a Machine Really Finish Your Sentence? (P19-1)

Copied to clipboard

Challenge: Existing commonsense models struggle to perform inferences that are trivial for humans, but are often misclassified by state-of-the-art models.
Approach: They propose a dataset that is adversarial to state-of-the-art commonsense reasoning and use it to build a model that is surprisingly robust.
Outcome: The proposed dataset is compared with existing models and scaled up towards a critical 'Goldilocks zone' wherein generated text is ridiculous to humans, yet often misclassified by state-of-the-art models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations