Papers with CommonsenseQA

33 papers
Token Level Routing Inference System for Edge Devices (2025.acl-demo)

Copied to clipboard

Challenge: Large language models (LLMs) have been gaining in performance but deployment in edge devices faces significant hurdles due to their high computational complexity.
Approach: They propose a collaborative decoding system that allows small models to perform on-device inference while selectively consulting a cloud-based large model for critical token generation.
Outcome: The proposed system achieves 60% performance gain on CommonsenseQA using a 0.5B model on an M1 MacBook, with under 7% of tokens generation uploaded to the large model in the cloud.
PipeNet: Question Answering with Semantic Pruning over Knowledge Graphs (2024.starsem-1)

Copied to clipboard

Challenge: Existing approaches to utilizing explicit knowledge graphs (KGs) are limited by the number of nodes in the subgraph.
Approach: They propose a grounding-pruning-reasoning pipeline to prune noisy nodes in subgraphs to improve the efficiency of graph reasoning with KG.
Outcome: The proposed method reduces computation cost and memory usage while obtaining decent representation of pruned subgraphs.
Common Sense or Ableism? Rethinking Commonsense Reasoning Through the Lens of Disability (2026.eacl-short)

Copied to clipboard

Challenge: a recent study finds that commonsense reasoning is not always universal and can leave disabled people behind . a case study of disabled people with long COVID shows that common sense is not universal .
Approach: They investigate how datasets and models deal with disability in commonsense reasoning . they use annotations from disabled and non-disabled persons for ableism .
Outcome: The proposed datasets have low sensitivity to human-detected ableism but still detect 5 to 25% of entries as ableist.
QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering (2021.naacl-main)

Copied to clipboard

Challenge: Existing question answering systems lack the ability to access relevant knowledge and reason over it.
Approach: They propose a model that uses KGs to identify relevant knowledge in QA contexts and perform joint reasoning over them.
Outcome: The proposed model improves on the CommonsenseQA and OpenBookQA datasets and performs interpretable and structured reasoning.
On Commonsense Cues in BERT for Solving Commonsense Tasks (2021.findings-acl)

Copied to clipboard

Challenge: Pre-trained language models can capture syntactic features, semantic information and factual knowledge, but structured commonsense knowledge is not captured well.
Approach: They quantitatively investigate the presence of structural commonsense cues in BERT when solving commonsensense tasks and the importance of such cue for the model prediction.
Outcome: The presence of commonsense knowledge is positively correlated to the model accuracy.
TSGP: Two-Stage Generative Prompting for Unsupervised Commonsense Question Answering (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies focus on acquiring relevant knowledge by retrieving external knowledge bases and fine-tuning pre-trained models.
Approach: They propose a two-stage prompt-based unsupervised commonsense question answering framework that leverages implicit knowledge stored in PrLMs to generate knowledge for questions with unlimited types and possible candidate answers independent of specified choices.
Outcome: The proposed framework significantly improves the reasoning ability of language models in unsupervised settings.
FROST: Factual Reasoning via Optimized Stochastic Trajectories in Large Language Models during Inference (2026.acl-industry)

Copied to clipboard

Challenge: Existing mitigation strategies are needed to improve large language models' reliability and efficiency.
Approach: They propose an inference-time framework that balances exploration andexploitation without additional training or context augmentation.
Outcome: FROST achieves 2–5 percentage point improvements over standard chain-of-thoughtprompting and reduces unsupported outputs by 40% relative to Standard CoT.
Generative Data Augmentation for Commonsense Reasoning (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in commonsense reasoning depend on large-scale human-authored training data.
Approach: They propose a generative data augmentation technique that augments human-authored training data by using pretrained language models.
Outcome: The proposed technique outperforms existing methods on commonsense reasoning benchmarks and enhances out-of-distribution generalization.
Scalable Multi-Hop Relational Reasoning for Knowledge-Aware Question Answering (2020.emnlp-main)

Copied to clipboard

Challenge: Existing work on augmenting question answering models with external knowledge (e.g., knowledge graphs) lacks transparency into the model’s prediction rationale.
Approach: They propose a knowledge-aware approach that equips pre-trained language models with a multi-hop relational reasoning module that performs multi-relational reasoning over subgraphs extracted from external knowledge graphs.
Outcome: The proposed model performs multi-hop, multi-relational reasoning over subgraphs extracted from external knowledge graphs.
Winnowing Knowledge for Multi-choice Question Answering (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing reasoning models suffer from noises in retrieved knowledge . encoding methods that use commonsense knowledge are less effective .
Approach: They propose a method which conducts interception and soft filtering to reduce noise . they use commonsense knowledge from Wikipedia and ConceptNet to encode questions and options .
Outcome: The proposed method improves on commonsense question answering tasks compared to baselines . it is able to conduct interception and soft filtering to shield the encoder from noise .
More Samples or More Prompts? Exploring Effective Few-Shot In-Context Learning for LLMs with In-Context Sampling (2024.findings-naacl)

Copied to clipboard

Challenge: Existing studies on LLM prompting focus on selecting a better set of data samples inside one single prompt input, but why not design and leverage multiple ICL prompts together to further improve the LLM’s performance?
Approach: They propose a low-resource LLM prompting technique to optimize the construction of multiple ICL prompt inputs to produce confident predictions.
Outcome: The proposed technique can produce confident predictions by optimizing the construction of multiple ICL prompt inputs on four NLI datasets and one QA dataset.
Dynamic Relevance Graph Network for Knowledge-Aware Question Answering (2022.coling-1)

Copied to clipboard

Challenge: Existing approaches to solve commonsense question answering problems often miss some edges between entities, which breaks the reasoning chain.
Approach: They propose a graph neural network architecture that uses relevance as graph edges to establish new edges dynamically for learning node representations in the graph network.
Outcome: The proposed approach shows competitive performance on two QA benchmarks, CommonsenseQA and OpenbookQA, compared to the state-of-the-art published results.
CORN: Co-Reasoning Network for Commonsense Question Answering (2022.coling-1)

Copied to clipboard

Challenge: Existing work uses two independent modules to model QA content and external commonsense knowledge graph (KG) Existing research uses two separate modules to create QA contextual text representations and relationships between QA entities.
Approach: They propose a commonsense question answering (QA) model that uses two independent modules to model QA contextual text representation and relationships between QA entities in KG.
Outcome: The proposed model achieves state-of-the-art on QA benchmarks in the CommonsenseQA and OpenBookQA datasets.
CommonGen: A Constrained Text Generation Challenge for Generative Commonsense Reasoning (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent studies show that pre-trained language models perform well on commonsense-reasoning benchmark datasets, but building machines with commonsence to compose plausible sentences remains challenging.
Approach: They propose a constrained text generation task for generative commonsense reasoning that generates a coherent sentence using common concepts.
Outcome: The proposed task generates a coherent sentence describing an everyday scenario using common concepts over 35k concept-sets.
Graph Reasoning for Question Answering with Triplet Retrieval (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to answer complex questions require reasoning over knowledge graphs (KGs) state-of-the-art methods constrain retrieved knowledge in local subgraphs and discard more diverse triplets that are disconnected but useful for question answering.
Approach: They propose a method to retrieve the most relevant triplets from KGs and then rerank them, which are then concatenated with questions to be fed into language models.
Outcome: The proposed method outperforms state-of-the-art methods on commonsenseQA and OpenbookQA datasets with 4.6% absolute accuracy.
I Know What You Asked: Graph Path Learning using AMR for Commonsense Reasoning (2020.coling-main)

Copied to clipboard

Challenge: a large amount of pre-defined commonsense knowledge is available for commonsensense reasoning . humans acquire commonsence in their lives, but machines cannot learn commonseense without assistance.
Approach: They propose an AMR-ConceptNet-Pruned (ACP) graph that is pruned from a full integrated graph . they show that the ACP graph interprets the reasoning path and predicts the correct answer .
Outcome: The proposed graph outperforms baseline models in the commonsenseQA task . it shows that the reasoning path can be interpreted with the relations and concepts provided by the graph .
Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training Data (2022.acl-long)

Copied to clipboard

Challenge: Experimental results show that REtrieving from the traINing datA only can lead to significant gains on multiple NLG and NLU tasks.
Approach: They propose to retrieve training instances from traINing datA and concatenate them with input to generate output.
Outcome: The proposed method achieves state-of-the-art results on XSum, BigPatent, and CommonsenseQA.
Explanations for CommonsenseQA: New Dataset and Models (2021.acl-long)

Copied to clipboard

Challenge: a dataset called CommonsenseQA (CQA) was recently released to advance the research on common-sense question answering (QA)
Approach: They propose to retrieve and generate explanations for a given question, correct answer choice, incorrect answer choices tuple from a dataset called CommonsenseQA.
Outcome: The proposed model beats baseline model by 100% in F1 score and similarity score of 61.9 .
KagNet: Knowledge-Aware Graph Networks for Commonsense Reasoning (D19-1)

Copied to clipboard

Challenge: empowering machines with the ability to perform commonsense reasoning has been seen as the bottleneck of artificial general intelligence .
Approach: They propose a textual inference framework that uses external commonsense knowledge graphs to answer commonsensical questions.
Outcome: The proposed framework is based on graph convolutional networks and LSTMs with a hierarchical path-based attention mechanism.
JointLK: Joint Reasoning with Language Models and Knowledge Graphs for Commonsense Question Answering (2022.naacl-main)

Copied to clipboard

Challenge: Existing KG-augmented models for commonsense question answering ignore the effectively fusing and reasoning over question context representations and the KG representations.
Approach: They propose a novel model which combines a logical reasoning and a dynamic pruning mechanism to solve these limitations.
Outcome: The proposed model improves existing models and performs interpretable reasoning on the CommonsenseQA and OpenBookQA datasets.
GraDA: Graph Generative Data Augmentation for Commonsense Reasoning (2022.coling-1)

Copied to clipboard

Challenge: Recent advances in commonsense reasoning have been fueled by the availability of large-scale human annotated datasets.
Approach: They propose a graph-generative data augmentation framework to synthesize factual data samples from knowledge graphs for commonsense reasoning.
Outcome: The proposed framework improves SocialIQA, CODAH, HellaSwag and CommonsenseQA . it also performs well for generative tasks like ProtoQA proving its robustness to adversaries .
Improving Unsupervised Commonsense Reasoning Using Knowledge-Enabled Natural Language Inference (2021.findings-emnlp)

Copied to clipboard

Challenge: Recent methods based on pre-trained language models have shown strong supervised performance on commonsense reasoning.
Approach: They propose to use a common framework to solve commonsense reasoning tasks using a dataset from NLI.
Outcome: The proposed method achieves state-of-the-art unsupervised performance on two commonsense reasoning tasks.
CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge (N19-1)

Copied to clipboard

Challenge: Recent work on question answering relies on factoid questions with little general knowledge.
Approach: They propose a dataset to capture commonsense question answering with prior knowledge . they extract multiple-choice questions that discriminate between the source and target concepts .
Outcome: The proposed dataset captures commonsense reasoning beyond associations . it obtains 56% accuracy, well below human performance, which is 89% .
Explaining and Improving Contrastive Decoding by Extrapolating the Probabilities of a Huge and Hypothetical LM (2024.emnlp-main)

Copied to clipboard

Challenge: Contrastive decoding (CD) improves the next-token distribution of a large expert language model (LM) using a small amateur LM.
Approach: They propose a new unsupervised decoding method called Asymptotic Probability Decoding (APD) that extrapolates the probability curves from the LMs of different sizes to infer the asymptototic probabilities from an infinitely large LM.
Outcome: The proposed method improves the next-token distribution of a large expert language model using a small amateur LM.
Explain Yourself! Leveraging Language Models for Commonsense Reasoning (P19-1)

Copied to clipboard

Challenge: Empirical results indicate that we can effectively leverage language models for commonsense reasoning.
Approach: They propose to use commonsense auto-generated explanations to train language models to generate explanations that can be used during training and inference in a commonsensense Auto-Generated Explanation framework.
Outcome: Empirical results show that the proposed framework improves on the commonsenseQA task by 10%.
ACENet: Attention Guided Commonsense Reasoning on Hybrid Knowledge Graph (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches estimate plausibility of candidate choices separately based on their respective KGs, without considering the interference among different choices.
Approach: They propose an Attention guided Commonsense rEasoning Network to integrate hybrid knowledge into the neural network.
Outcome: The proposed model outperforms existing methods on CommonsenseQA and OpenbookQA datasets and shows significant performance gains.
Explanation Graph Generation via Generative Pre-training over Synthetic Graphs (2023.findings-acl)

Copied to clipboard

Challenge: Existing frameworks for explanation graph generation are limited due to the large number of datasets available.
Approach: They propose a text-to-graph generative task to pre-train a model to bridge the text-graph gap.
Outcome: The proposed framework surpasses all baseline systems with remarkable margins on ExplaGraphs and CommonsenseQA.
FORK: A Bite-Sized Test Set for Probing Culinary Cultural Biases in Commonsense Reasoning Models (2023.findings-acl)

Copied to clipboard

Challenge: a recent study shows that commonsense knowledge is universally shared by most people . early efforts to schematize commonsensical knowledge as scripts provide examples of unintended biases .
Approach: They propose a set of questions for probing cultural biases and assumptions in commonsense reasoning systems . they test commonsensibleQA-style questions on food-related customs in the u.s.
Outcome: The proposed questions show that they are better at detecting biases in commonsense reasoning systems than on non-US cultures.
Dynamic Heterogeneous-Graph Reasoning with Language Models and Knowledge Representation Learning for Commonsense Question Answering (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for QA use knowledge graphs, but they ignore subgraph optimization and subgraph deepening.
Approach: They propose a dynamic heterogeneous-graph reasoning method with LMs and knowledge representation learning that optimizes the structure and knowledge representing of the HKG using a two-stage pruning strategy and knowledge-representation learning.
Outcome: The proposed method improves on existing methods at CommonsenseQA and OpenBookQA.
VarBench: Robust Language Model Benchmarking Through Dynamic Variable Perturbation (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent benchmarks release only training and validation sets, keeping the test set labels closed-source.
Approach: They propose to extract variables from each test case and define a value range for each variable.
Outcome: The proposed method improves the accuracy of the evaluations on four datasets covering mathematical generation and multiple-choice tasks.
FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing models generate explanations that appear coherent while containing unfaithful intermediate steps.
Approach: They propose a causality-inspired framework for evaluating CoT quality using controlled perturbations as an instrumental signal to separate genuine step-to-step dependence from bias-driven artifacts.
Outcome: Experiments on GSM8K, MATH, and CommonsenseQA show that FACT-E improves reasoning-trajectory selection and yields stronger in-context learning exemplars.
Towards Faithful Knowledge Graph Explanation Through Deep Alignment in Commonsense Question Answering (2024.emnlp-main)

Copied to clipboard

Challenge: Current methods for generating faithful explanations overlook path decoding faithfulness, leading to divergence between graph encoder outputs and model predictions.
Approach: They propose an algorithm to assess KG representation reliability and an LM-KG distribution-aware Alignment algorithm to improve explanation faithfulness without ground truth.
Outcome: The proposed algorithm improves explanation faithfulness without ground truth and significantly improves fidelity and model performance.
Mixing Inference-time Experts for Enhancing LLM Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for improving reasoning quality in large language models are limited to using a single expert.
Approach: They propose a framework that finetunes and merges expert logits from one LLM . they use commonsense and entailment reasoning experts to improve chain-of-thought reasoning .
Outcome: The proposed framework outperforms baselines on three question-answering datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations