Challenge: Existing benchmarks measure common sense knowledge indirectly or without reasoning.
Approach: They propose a benchmark to test whether a system can differentiate natural language statements that make sense from those that do not make sense.
Outcome: The proposed benchmarks show that models trained on large corpora perform better than humans on some benchmarks.

Similar Papers

Commonsense Reasoning for Natural Language Processing (2020.acl-tutorials)

Copied to clipboard

Challenge: In this tutorial, we will outline the various types of commonsense knowledge and discuss techniques to gather and represent commonsence knowledge.
Approach: This tutorial will provide researchers with the critical foundations and recent advances in commonsense representation and reasoning.
Outcome: This tutorial will outline the various types of commonsense and discuss techniques to gather and represent commonsence knowledge while highlighting the challenges specific to this type of knowledge (e.g., reporting bias).
What Will it Take to Fix Benchmarking in Natural Language Understanding? (2021.naacl-main)

Copied to clipboard

Challenge: Evaluation for many natural language understanding (NLU) tasks is broken due to unreliable and biased systems scoring so high on standard benchmarks.
Approach: They argue that current benchmarks fail at four criteria for evaluation . they argue that adversarial data collection does not address the causes of failures .
Outcome: The proposed frameworks fail at four criteria, and adversarial data collection does not address the causes of these failures, the authors argue . restoring a healthy evaluation ecosystem will require significant progress in the design of benchmark datasets, reliability with which they are annotated, their size, and the ways they handle social bias.
On General Language Understanding (2023.findings-emnlp)

Copied to clipboard

Challenge: a recent paper suggests that the evidence underspecifies the understanding of large language models.
Approach: They propose to use a "general language understanding" benchmark to examine what it could mean in machines.
Outcome: The proposed model can be used to ground questions of the adequacy of benchmarking methods.
Curious Case of Language Generation Evaluation Metrics: A Cautionary Tale (2020.coling-main)

Copied to clipboard

Challenge: a few popular metrics are still used to evaluate language generation systems despite their known limitations.
Approach: They propose to use automatic metrics to evaluate language generation systems . they show that they prefer system outputs to human-authored texts .
Outcome: The proposed metrics are insensitive to correct translations of rare words and can yield high scores when given a single sentence as system output for the entire test set.
Nibbling at the Hard Core of Word Sense Disambiguation (2022.acl-long)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a task that is based on a set of pre-trained language models.
Approach: They propose to use Word Sense Disambiguation to test whether systems can handle ambiguous words.
Outcome: The proposed benchmarks show that seven of the most representative state-of-the-art systems make trivial errors on traditional evaluation benchmarks.
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing work investigating the logical reasoning ability of large language models has focused only on a couple of inference rules of propositional and first-order logics.
Approach: They propose to use a natural language question-answering dataset to evaluate the logical reasoning ability of large language models.
Outcome: The proposed model performs poorly on a range of natural language questions using chain-of-thought prompting.
Reasoning about Uncertainty: Do Reasoning Models Know When They Don’t Know? (2026.findings-eacl)

Copied to clipboard

Challenge: Reasoning models are prone to generating confident, plausible responses that are incorrect (hallucinations).
Approach: They introduce introspective uncertainty quantification to examine whether reasoning models are well-calibrated and does deeper reasoning improve their calibration?
Outcome: The proposed model calibrations show that models are overconfident, overconfent and overconfust with deeper reasoning.
Which Evaluations Uncover Sense Representations that Actually Make Sense? (2020.lrec-1)

Copied to clipboard

Challenge: Existing sense representations fail for human-centric tasks like inspecting a language’s sense inventory.
Approach: They propose a coherence evaluation for sense embeddings and a model optimized for finding interpretable sense representations that are more coherent than existing sense embeds.
Outcome: The proposed model is more coherent than existing sense embeddings and offers comparable word similarities with multisense representations while learning more distinguishable, interpretable senses.
How Pre-trained Word Representations Capture Commonsense Physical Comparisons (D19-60)

Copied to clipboard

Challenge: Pre-trained word representations capture common sense on physical properties such as size and weight.
Approach: They investigate whether pre-trained representations capture comparisons and find they have higher accuracy than previous approaches.
Outcome: The proposed models learn a consistent ordering over all the objects in the comparisons.
A Pragmatics-Centered Evaluation Framework for Natural Language Understanding (2022.lrec-1)

Copied to clipboard

Challenge: a number of studies have suggested that models induce universal text representations . current benchmarks focus on semantic phenomena, so pragmatics needs to be the focus .
Approach: They propose a benchmark that unites 11 pragmatics-focused evaluation datasets for English.
Outcome: The proposed benchmark shows that natural language inference does not result in genuinely universal representations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations