How well do NLI models capture verb veridicality? (D19-1)

Copied to clipboard

Challenge: In natural language inference, contexts are considered veridical if they allow us to infer that their underlying propositions make true claims about the real world.
Approach: They propose to use a dataset for veridicality evaluation consisting of 1,500 sentence pairs, covering 137 unique verbs.
Outcome: The proposed model learns to make correct inferences about veridicality in verb-complement constructions.

Similar Papers

Exploring Transitivity in Neural NLI Models through Veridicality (2021.eacl-main)

Copied to clipboard

Challenge: Despite recent success of deep neural networks in natural language processing, the extent to which they can demonstrate human-like generalization capacities remains unclear.
Approach: They propose an analysis method to evaluate whether models can draw inferences composed of veridical inference and arbitrary inference types.
Outcome: The proposed model performs poorly on transitivity inference tasks, suggesting it lacks generalization capacity for drawing composite inferences from training examples.
Are Natural Language Inference Models IMPPRESsive? Learning IMPlicature and PRESupposition (2020.acl-main)

Copied to clipboard

Challenge: Natural language inference (NLI) is an increasingly important task for natural language understanding . however, the ability of NLI models to make pragmatic inferences remains understudied .
Approach: They use semi-automatically generated sentence pairs to evaluate whether NLI models make pragmatic inferences.
Outcome: The proposed model trains on multiNLI and shows that it learns to draw pragmatic inferences.
How Fast can BERT Learn Simple Natural Language Inference? (2021.eacl-main)

Copied to clipboard

Challenge: Efficiency of learning of BERT is very slow due to hidden dataset bias . however, some studies show that it can learn with surface clues/patterns .
Approach: They propose to use a simple entailment judgment case to test whether BERT can learn without hidden dataset bias.
Outcome: The proposed case shows that BERT can learn without hidden bias without utilizing dataset bias.
Stretching Sentence-pair NLI Models to Reason over Long Documents and Clusters (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in modeling and datasets demonstrate promising performance for NLI.
Approach: They explore the direct zero-shot applicability of NLI models to real applications . they analyze the robustness of models to longer and out-of-domain inputs .
Outcome: The proposed models are robust to longer and out-of-domain inputs and can perform on full documents.
Simple but Challenging: Natural Language Inference Models Fail on Simple Sentences (2022.findings-emnlp)

Copied to clipboard

Challenge: Natural language inference (NLI) tasks are difficult to perform on large datasets . a small number of simple sentences can improve model performance, authors say .
Approach: They propose to use syntactically simple sentences to test the inference ability of NLI models.
Outcome: The proposed set of simple sentences shows that the models fine-tuned on MNLI and SNLI perform poorly on Simple Pair.
Evaluating BERT for natural language inference: A case study on the CommitmentBank (D19-1)

Copied to clipboard

Challenge: Natural language inference datasets can identify premise-hypothesis relationship without observing premise . recasting of the CommitmentBank for NLI creates hypotheses that stand in entailment/contradiction/neutral relationship with premise.
Approach: They propose to recast the CommitmentBank for NLI to stand in certain relationships with the premise . hypotheses are complements of clause-embedding verbs in each premise, rethinking the CommittedBank .
Outcome: The proposed model performs well on the CommitmentBank with 85% F1 . however, the model does not capture the full complexity of pragmatic reasoning, authors say .
Breaking NLI Systems with Sentences that Require Simple Lexical Inferences (P18-2)

Copied to clipboard

Challenge: a new test set shows the deficiency of state-of-the-art models in inferences that require lexical and world knowledge.
Approach: They create a new NLI test set that shows the deficiency of state-of-the-art models in inferences that require lexical and world knowledge.
Outcome: The new examples are simpler than the SNLI test set, but the state-of-the-art systems perform poorly on it.
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: In the recent past, a popular way of evaluating natural language understanding was to consider a model’s ability to perform natural language inference (NLI) tasks.
Approach: They focus on five different NLI benchmarks across six models of different scales and examine how their accuracies develop during training.
Outcome: The softmax distributions of models align with human label distributions in cases where statements are ambiguous or vague.
Investigating representations of verb bias in neural language models (2020.emnlp-main)

Copied to clipboard

Challenge: Languages typically provide more than one grammatical construction to express certain types of messages.
Approach: They propose a large benchmark dataset containing 50K human judgments for 5K distinct sentence pairs in the English dative alternation.
Outcome: The proposed model outperforms recurrent architectures even under comparable parameter and training settings.
End-to-End Bias Mitigation by Modelling Biases in Corpora (2020.acl-main)

Copied to clipboard

Challenge: Recent studies have shown that strong natural language understanding models are prone to relying on unwanted dataset biases without learning the underlying task.
Approach: They propose two learning strategies to train neural models that are more robust to dataset biases and transfer better to out-of-domain datasets.
Outcome: The proposed methods improve robustness in all settings and transfer better to out-of-domain datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations