DIVE: Towards Descriptive and Diverse Visual Commonsense Generation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Towards human-level visual understanding, visual commonsense generation has been introduced . but current research on visual commonense generation ignores an important human cognitive ability . |
| Approach: | They propose a visual commonsense generation framework to improve inferences by visual common sense generation. |
| Outcome: | The proposed framework outperforms state-of-the-art models in descriptiveness and diversity . human evaluations confirm that the framework aligns closely with human judgments on descriptiveness . |
Similar Papers
VCD: A Dataset for Visual Commonsense Discovery in Images (2025.findings-acl)
Copied to clipboard
| Challenge: | Visual commonsense data sets lack visual grounded representations of commonsensense . existing knowledge bases lack visual-based knowledge tied to actual visual scenes . |
| Approach: | They present a large-scale visual commonsense dataset with over 100,000 images and 14 million object-commonsense pairs that integrates both Seen (directly observable) and Unseen (inferrable) commonsens. |
| Outcome: | The proposed model integrates Seen (directly observable) and Unseen (inferrable) commonsense across Property, Action, and Space aspects. |
Think Beyond Words: Exploring Context-Relevant Visual Commonsense for Diverse Dialogue Generation (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to generate intelligent open-domain dialogue agents only consider auxiliary commonsense stored in pure text, ignoring grounding information from the external visual world. |
| Approach: | They propose a VIsual Commonsense enhanced dialogue generaTOR that exploits auxiliary commonsense from images related to context to generate coherent and informative responses. |
| Outcome: | The proposed method outperforms the latest competitive methods in terms of coherence and diversity on two public datasets. |
Improving Diversity of Commonsense Generation by Large Language Models via In-Context Learning (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown proficiency in enhancing the generation quality across various tasks without the need for any fine-tuning. |
| Approach: | They propose a method that diversifies the LLM generations while preserving their quality. |
| Outcome: | The proposed method can be used as training data to improve diversity in existing commonsense generators. |
Commonsense Reasoning for Natural Language Processing (2020.acl-tutorials)
Copied to clipboard
| Challenge: | In this tutorial, we will outline the various types of commonsense knowledge and discuss techniques to gather and represent commonsence knowledge. |
| Approach: | This tutorial will provide researchers with the critical foundations and recent advances in commonsense representation and reasoning. |
| Outcome: | This tutorial will outline the various types of commonsense and discuss techniques to gather and represent commonsence knowledge while highlighting the challenges specific to this type of knowledge (e.g., reporting bias). |
Improving Unsupervised Commonsense Reasoning Using Knowledge-Enabled Natural Language Inference (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Recent methods based on pre-trained language models have shown strong supervised performance on commonsense reasoning. |
| Approach: | They propose to use a common framework to solve commonsense reasoning tasks using a dataset from NLI. |
| Outcome: | The proposed method achieves state-of-the-art unsupervised performance on two commonsense reasoning tasks. |
Evaluating the Evaluation of Diversity in Commonsense Generation (2025.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation metrics for commonsense generation are unclear on which metrics are best suited for evaluating the diversity of outputs. |
| Approach: | They propose to use a large language model to analyze commonsense generation data to determine which diversity metrics are best suited for commonsensing. |
| Outcome: | The proposed metrics outperform form-based metrics and show high correlations with the LLM-based ratings. |
Divergent Thinking: Escape the Homogeneity Trap in Generative Commonsense Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Generative commonsense reasoning requires models to synthesize coherent narratives that satisfy lexical constraints and commonsensical logic. |
| Approach: | They propose a framework that allows for deep semantic diversity rather than surface-level lexical variation. |
| Outcome: | The proposed framework achieves over 10% improvement in overall accuracy on NoRa and SPICE score on CommonGen-Lite. |
Find Someone Who: Visual Commonsense Understanding in Human-Centric Grounding (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Visual scenes often involve multiple people and humans can distinguish between them based on context descriptions about what happened before, their mental/physical states, and intentions. |
| Approach: | They propose a task that tests human-centric commonsense grounding models' ability to distinguish individuals given context descriptions about what happened before and their mental/physical states or intentions. |
| Outcome: | The proposed model outperforms pre-trained and non-pretrained models on 130k commonsense descriptions annotated on 67k images. |
Synthetic Data Generation for Training Diversified Commonsense Reasoning Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing Generative Commonsense Reasoning datasets are created using a small number of human annotators, covering only a narrow set of commonsense scenarios. |
| Approach: | They propose to use a synthetic dataset to train diverse commonsense generators. |
| Outcome: | The proposed model improves both generation diversity and quality compared with vanilla models and human-crafted datasets across different size Large Language Models (LLMs). |
Video2Commonsense: Generating Commonsense Descriptions to Enrich Video Captioning (2020.emnlp-main)
Copied to clipboard
| Challenge: | Observable changes in the scene are reflected in captions, but actions are also linked to social aspects such as intentions, effects, and attributes that describe the agent. |
| Approach: | They propose to generate captions from videos that describe latent aspects of the human agent's actions. |
| Outcome: | The proposed model can be used to describe latent aspects of human actions in video clips and answer questions about videos. |