Misleading Failures of Partial-input Baselines (P19-1)

Copied to clipboard

Challenge: Recent work establishes dataset difficulty and removes annotation artifacts via partial-input baselines.
Approach: They propose to use partial-input baselines to establish dataset difficulty . they show how trivial patterns only visible in the full input can evade partial-output baseline .
Outcome: The proposed model can solve 15% of previously-thought "hard" examples.

Similar Papers

Partial-input baselines show that NLI models can ignore context, but they don’t. (2022.naacl-main)

Copied to clipboard

Challenge: Researchers have shown that many datasets contain statistical biases, or "annotation artifacts" that systems leverage to correctly predict entailment.
Approach: They propose to use edited contexts to examine RoBERTa models' sensitivity to edited context to examine their model's sensitivity.
Outcome: The proposed model can learn to condition on context, despite being trained on artifact-ridden datasets.
Even the Simplest Baseline Needs Careful Re-investigation: A Case Study on XML-CNN (2022.naacl-main)

Copied to clipboard

Challenge: XML-CNN has been a popular research topic in NLP due to its superior performance . however, the increasing complexity brings difficulties to ensure the true architectural progress .
Approach: They propose to re-examine an influential multi-label text classification method . they propose suitable baselines for multi-level text classification tasks .
Outcome: The proposed method performs better than the original model, the authors show . they show that the re-implementation reveals contradictory results to the original work .
Breaking NLI Systems with Sentences that Require Simple Lexical Inferences (P18-2)

Copied to clipboard

Challenge: a new test set shows the deficiency of state-of-the-art models in inferences that require lexical and world knowledge.
Approach: They create a new NLI test set that shows the deficiency of state-of-the-art models in inferences that require lexical and world knowledge.
Outcome: The new examples are simpler than the SNLI test set, but the state-of-the-art systems perform poorly on it.
Revisiting Pathologies of Neural Models under Input Reduction (2023.findings-acl)

Copied to clipboard

Challenge: Recent studies have shown that modern neural models tend to be miscalibrated.
Approach: They examine why models produce high-confidence predictions on inputs that appear nonsensical to humans . previous work suggested that models fail to assign low probabilities due to model overconfidence .
Outcome: The proposed methods can be extended to reduce the number of examples but with the cost of miscalibration.
A guide to the dataset explosion in QA, NLI, and commonsense reasoning (2020.coling-tutorials)

Copied to clipboard

Challenge: a tutorial aims to provide an up-to-date guide to the recent datasets . the target audience is the NLP practitioners who are lost in dozens of the recent data sets.
Approach: This tutorial provides an up-to-date guide to the recent datasets . it surveys old and new methodological issues with dataset construction .
Outcome: This tutorial aims to provide an up-to-date guide to the recent datasets . it surveys the old and new methodological issues with dataset construction .
Strong Baselines for Simple Question Answering over Knowledge Graphs with and without Neural Networks (N18-2)

Copied to clipboard

Challenge: Existing work on simple question answering over knowledge graphs involves increasingly complex NN architectures.
Approach: They propose to decompose the problem into entity detection, entity linking, relation prediction, evidence combination and heuristics.
Outcome: The proposed approach outperforms existing models and benchmarks on a simple QA task.
MedNLI Is Not Immune: Natural Language Inference Artifacts in the Clinical Domain (2021.acl-short)

Copied to clipboard

Challenge: a large number of crowdworker-constructed datasets have been used to conduct natural language inference (NLI) on unstructured, domainspecific texts such as patient notes, pathology reports, and scientific papers.
Approach: They investigate whether MedNLI contains lexical and syntactic annotation artifacts associated with annotation process that allow hypothesis-only classifiers to achieve better-than-random performance.
Outcome: The proposed model outperforms a majority-class baseline model on a physician-annotated dataset with premises extracted from clinical notes.
Are Neural Topic Models Broken? (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation paradigms are often divorced from real-world use . recent results have challenged the validity of the prevailing model evaluation paradigm .
Approach: They show that neural topic models fare worse in both respects compared to an established classical method.
Outcome: The proposed method outperforms the members of the ensemble in both respects.
An Empirical Study of Building a Strong Baseline for Constituency Parsing (P18-2)

Copied to clipboard

Challenge: Sequence-to-sequence models have been used for natural language generation tasks such as machine translation and summarization.
Approach: They propose to build a strong baseline based on general purpose sequence-to-sequence models for constituency parsing.
Outcome: The proposed model outperforms existing models in natural language generation tasks without any explicit task-specific knowledge or architecture of constituent parsing.
On the Transferability of Minimal Prediction Preserving Inputs in Question Answering (2021.naacl-main)

Copied to clipboard

Challenge: Recent work establishes the presence of short, uninterpretable input fragments that yield high confidence and accuracy in neural models.
Approach: They investigate competing hypotheses for the existence of MPPIs in question answering . they discover a perplexing invariance of MPIs to random training seed, model architecture, pretraining, and training domain.
Outcome: The proposed model performance is higher than comparable short queries.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations