| Challenge: | a dataset for visual reasoning with natural language and images is available. |
| Approach: | They propose a dataset for joint reasoning about natural language and images . they crowdsource 107,292 examples of English sentences paired with web photographs . |
| Outcome: | The proposed dataset combines 107,292 examples of English sentences with web photographs . Qualitative analysis shows the data requires compositional joint reasoning . |
Similar Papers
Visually Grounded Reasoning across Languages and Cultures (2021.emnlp-main)
Copied to clipboard
| Challenge: | a new protocol allows for a multilingual hierarchy of concepts and images based on native speakers . the results suggest that the current models are not robust enough to handle multilingual data . |
| Approach: | They propose a protocol to construct an ImageNet-style hierarchy representative of more languages and cultures. |
| Outcome: | The proposed protocol lets the selection of concepts and images be entirely driven by native speakers, rather than scraping them automatically. |
Natural Language Rationales with Full-Stack Visual Reasoning: From Pixels to Semantic Frames to Commonsense Graphs (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models that use natural language rationales provide intuitive, higher-level explanations that are easily understandable by humans. |
| Approach: | They propose a model that generates free-text rationales by combining pretrained language models with object recognition, grounded visual semantic frames, and visual commonsense graphs. |
| Outcome: | The proposed model generates free-text rationales by combining pretrained language models with object recognition, grounded visual semantic frames, and visual commonsense graphs. |
Pushing the Limits of Radiology with Joint Modeling of Visual and Textual Information (P18-3)
Copied to clipboard
| Challenge: | Recent research has focused on the intersection of computer vision and natural language processing, but its adaption to the medical domain is not fully explored. |
| Approach: | They aim to develop machine learning models that can reason jointly on medical images and clinical text for advanced search, retrieval, annotation and description of medical images. |
| Outcome: | The proposed models can reason jointly on medical images and clinical text for advanced search, retrieval, annotation and description of medical images. |
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference (N18-1)
Copied to clipboard
| Challenge: | et al., 1996, show that many of the most actively studied problems in NLP depend in large part on natural language understanding (NLU). |
| Approach: | They propose a dataset for machine learning that uses ten different genres of English to evaluate sentences for their meanings. |
| Outcome: | The multi-genre natural language inference corpus is one of the largest available for natural language understanding. |
Neural Naturalist: Generating Fine-Grained Image Comparisons (D19-1)
Copied to clipboard
| Challenge: | a dataset of 41k sentences describes fine-grained differences between photographs of birds . human observers are adept at making fine-grain comparisons, but sometimes require aid in distinguishing visually similar classes. |
| Approach: | They propose a model that generates comparative language from a dataset of 41k sentences describing fine-grained differences between photographs of birds. |
| Outcome: | The proposed model can explain differences in visual embedding space using natural language . it evaluates the results with humans who must use the descriptions to distinguish real images . |
Representing Verbs with Visual Argument Vectors (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing models for verb semantic similarities are based on linguistic data, but they do not register intuitive attributes. |
| Approach: | They evaluated two textual distributional semantic models and a visual one to explore verb semantic similarities. |
| Outcome: | The proposed models extract meaningful information and capture semantic similarity between verbs using visual distributional models. |
CITE: A Corpus of Image-Text Discourse Relations (N19-1)
Copied to clipboard
| Challenge: | a crowd-sourced resource characterizes inferences in image-text contexts in the domain of cooking recipes . a recent study has found that image-image presentations are more effective at integrating text and image . |
| Approach: | They propose a crowd-sourced resource for multimodal discourse characterizing inferences in image-text contexts in the domain of cooking recipes in the form of coherence relations. |
| Outcome: | The proposed corpus enables a better understanding of communication and common-sense reasoning . it is particularly important for automating the understanding and generation of text-image presentations . |
A Probabilistic Model for Joint Learning of Word Embeddings from Texts and Images (D18-1)
Copied to clipboard
| Challenge: | Existing approaches combine language and perception to infer word embeddings . however, the embeddables produced by such models do not reflect the actual word representations. |
| Approach: | They propose a probabilistic model that integrates linguistic and perceptual inputs to explain observed word-context pairs in a text corpus. |
| Outcome: | The proposed model achieves competitive or stronger results on tasks of assessing pairwise word similarity and image/caption retrieval compared to other state-of-the-art models. |
Collecting Diverse Natural Language Inference Problems for Sentence Representation Evaluation (D18-1)
Copied to clipboard
Adam Poliak, Aparajita Haldar, Rachel Rudinger, J. Edward Hu, Ellie Pavlick, Aaron Steven White, Benjamin Van Durme
| Challenge: | a plethora of new natural language inference datasets has been created in recent years . however, these datasets do not provide clear insight into what type of reasoning or inference a model may be performing. |
| Approach: | They propose to recast 13 existing natural language inference datasets into a common structure. |
| Outcome: | The proposed datasets provide insight into how well a sentence representation captures distinct types of reasoning. |
Probing Image-Language Transformers for Verb Understanding (2021.findings-acl)
Copied to clipboard
| Challenge: | Multimodal image-language transformers have achieved impressive results on a variety of tasks that rely on fine-tuning. |
| Approach: | They collect a dataset of image-sentence pairs consisting of 421 verbs . they evaluate pretrained image-language transformers and find they fail more in situations that require verb understanding compared to other parts of speech. |
| Outcome: | The proposed model trains on a manually-annotated and smaller dataset does better on the task. |