Challenge: Videos often capture objects, their visible properties, their motion, and the interactions between different objects.
Approach: They propose a video question answering dataset for reasoning about the implicit physical properties of objects in a scene.
Outcome: The proposed dataset enables evaluation under several out-of-distribution settings – videos with objects with masses, coefficients of friction, and initial velocities that are not observed in the training distribution.

Similar Papers

ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos (2023.emnlp-main)

Copied to clipboard

Challenge: despite its importance, there are few datasets that cover multimodal counterfactual reasoning . a dataset focusing on this area is limited because of its limited coverage over synthetic environments .
Approach: They develop a video question answering dataset that provides questions on multimodal reasoning . they ask questions about counterfactual hypotheses over visual events .
Outcome: The proposed dataset shows a significant performance gap between models and humans . it provides questions that span physical, social, and temporal dimensions .
IfQA: A Dataset for Open-domain Question Answering under Counterfactual Presuppositions (2023.emnlp-main)

Copied to clipboard

Challenge: Existing open-domain QA tasks focus on questions whose answer can be deduced directly from global factual knowledge.
Approach: They propose a dataset where each question is based on a counterfactual presupposition via an "if" clause.
Outcome: The IfQA dataset contains 3,800 questions that were annotated by crowdworkers on relevant Wikipedia passages.
CommVQA: Situating Visual Question Answering in Communicative Contexts (2024.emnlp-main)

Copied to clipboard

Challenge: Current visual question answering models are trained on image-question pairs in isolation, but the questions people ask are dependent on their informational needs and prior knowledge about the image content.
Approach: They propose a visual question-answer-as-question dataset that contains 1000 images and 8,949 question-announcer pairs to evaluate how situating images within naturalistic contexts shapes visual questions.
Outcome: The proposed dataset contains 1000 images and 8,949 question-answer pairs.
Learning to Contrast the Counterfactual Samples for Robust Visual Question Answering (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods of generating counterfactual samples are not fully utilized in the task of Visual Question Answering (VQA).
Approach: They propose a self-supervised contrastive learning mechanism to learn the relationship between original samples, factual samples and counterfactual samples.
Outcome: The proposed method surpasses state-of-the-art models on the VQA-CP dataset, a diagnostic benchmark for assessing the VQ model’s robustness.
‘Just because you are right, doesn’t mean I am wrong’: Overcoming a bottleneck in development and evaluation of Open-Ended VQA tasks (2021.eacl-main)

Copied to clipboard

Challenge: Existing visual question answering datasets assume only one ground truth answer for each question.
Approach: They propose alternative answer sets (AAS) of ground-truth answers to address this limitation . they modify top VQA solvers to support multiple plausible answers for a question .
Outcome: The proposed approach improves on the GQA dataset and shows that it is more efficient than previous approaches.
Discovering the Unknown Knowns: Turning Implicit Knowledge in the Dataset into Explicit Training Examples for Visual Question Answering (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods address this issue by introducing an auxiliary task such as visual grounding, cycle consistency, or debiasing.
Approach: They propose a data augmentation pipeline to turn “known” knowledge into training examples for VQA.
Outcome: The proposed model can handle multi-modal information and is based on human-annotated examples.
LifeQA: A Real-life Dataset for Video Question Answering (2020.lrec-1)

Copied to clipboard

Challenge: Existing video question answering datasets consist of movies and TV shows, but they are not representative of our day-to-day lives.
Approach: They propose a benchmark dataset for video question answering that focuses on day-to-day situations.
Outcome: The proposed dataset analyzes the challenging but realistic aspects of LifeQA . it consists of video clips and over 2.3k multiple-choice questions .
PROST: Physical Reasoning about Objects through Space and Time (2021.findings-acl)

Copied to clipboard

Challenge: a new dataset is available to test pretraining of physical reasoning models . state-of-the-art models are inadequate at reasoning about physical interactions, authors say .
Approach: They present a dataset that contains 18,736 multiple-choice questions from 14 templates . they propose to use the dataset to probe both causal and masked language models .
Outcome: The proposed dataset contains 18,736 multiple-choice questions covering 10 physical reasoning concepts.
How Pre-trained Word Representations Capture Commonsense Physical Comparisons (D19-60)

Copied to clipboard

Challenge: Pre-trained word representations capture common sense on physical properties such as size and weight.
Approach: They investigate whether pre-trained representations capture comparisons and find they have higher accuracy than previous approaches.
Outcome: The proposed models learn a consistent ordering over all the objects in the comparisons.
CRASS: A Novel Data Set and Benchmark to Test Counterfactual Reasoning of Large Language Models (2022.lrec-1)

Copied to clipboard

Challenge: CRASS data set and benchmark provide novel test scheme to evaluate large language models . authors present and explain the CRAS data set, a novel basis to test reasoning and natural language understanding of LLMs .
Approach: They introduce a new test scheme utilizing questionized counterfactual conditionals to evaluate large language models.
Outcome: The proposed model sets out to be the most powerful and valid tool to evaluate large language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations