Hoa Trong Vu, Claudio Greco, Aliia Erofeeva, Somayeh Jafaritazehjan, Guido Linders, Marc Tanti, Alberto Testoni, Raffaella Bernardi, Albert Gatt
| Challenge: | Existing models for entailment analysis are not performing well in visual information-based models. |
| Approach: | They propose to use a visual representation of the Textual Entailment task to compare visual-grounded models with a multimodal version of the SNLI dataset. |
| Outcome: | The proposed model performs better when there is an image of the “world” or “situation” . |
Similar Papers
Multimodal Logical Inference System for Visual-Textual Entailment (P19-2)
Copied to clipboard
| Challenge: | Recent studies of multimodal inference provide challenging tasks such as visual question answering and visual reasoning. |
| Approach: | They propose an unsupervised multimodal logical inference system that can prove entailment relations between texts and images by combing semantic parsing and theorem proving. |
| Outcome: | The proposed system can handle semantically complex sentences for visual-textual inference. |
Visually Grounded Reasoning across Languages and Cultures (2021.emnlp-main)
Copied to clipboard
| Challenge: | a new protocol allows for a multilingual hierarchy of concepts and images based on native speakers . the results suggest that the current models are not robust enough to handle multilingual data . |
| Approach: | They propose a protocol to construct an ImageNet-style hierarchy representative of more languages and cultures. |
| Outcome: | The proposed protocol lets the selection of concepts and images be entirely driven by native speakers, rather than scraping them automatically. |
A Corpus for Reasoning about Natural Language Grounded in Photographs (P19-1)
Copied to clipboard
| Challenge: | a dataset for visual reasoning with natural language and images is available. |
| Approach: | They propose a dataset for joint reasoning about natural language and images . they crowdsource 107,292 examples of English sentences paired with web photographs . |
| Outcome: | The proposed dataset combines 107,292 examples of English sentences with web photographs . Qualitative analysis shows the data requires compositional joint reasoning . |
It is not a piece of cake for GPT: Explaining Textual Entailment Recognition in the presence of Figurative Language (2025.coling-main)
Copied to clipboard
| Challenge: | Figure-based language is used to convey opinions, ideas, or emotions in texts and dialogues. |
| Approach: | They evaluate the capabilities of Large Language Models to address TER and generate textual explanations of TER predictions. |
| Outcome: | The proposed model outperforms the open-source models in Zero- and Few-Shot Learning settings and shows significant performance improvements. |
Understanding Figurative Meaning through Explainable Visual Entailment (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing models for visual entailment and visual question-answering have limited ability to understand figurative meaning in images and captions. |
| Approach: | They propose a task framing the figurative meaning understanding problem as an explainable visual entailment task where the model has to predict whether the image entitles a caption and justify the predicted label with a textual explanation. |
| Outcome: | The proposed dataset contains 6,027 image, caption, label, explanation instances covering five diverse figurative phenomena. |
A Probabilistic Model for Joint Learning of Word Embeddings from Texts and Images (D18-1)
Copied to clipboard
| Challenge: | Existing approaches combine language and perception to infer word embeddings . however, the embeddables produced by such models do not reflect the actual word representations. |
| Approach: | They propose a probabilistic model that integrates linguistic and perceptual inputs to explain observed word-context pairs in a text corpus. |
| Outcome: | The proposed model achieves competitive or stronger results on tasks of assessing pairwise word similarity and image/caption retrieval compared to other state-of-the-art models. |
Visual-Textual Entailment with Quantities Using Model Checking and Knowledge Injection (2024.lrec-main)
Copied to clipboard
| Challenge: | Visual-textual entailment (VTE) is a critical task in multimodal inference. |
| Approach: | They propose a visual-textual entailment system that solves VTE tasks with quantities and negation. |
| Outcome: | The proposed system solves visual-textual entailment tasks with quantities and negation more robustly than previous approaches. |
Premise-based Multimodal Reasoning: Conditional Inference on Joint Textual and Visual Clues (2022.acl-long)
Copied to clipboard
Qingxiu Dong, Ziwei Qin, Heming Xia, Tian Feng, Shoujie Tong, Haoran Meng, Lin Xu, Zhongyu Wei, Weidong Zhan, Baobao Chang, Sujian Li, Tianyu Liu, Zhifang Sui
| Challenge: | Existing work in vision language cross-modal reasoning uses binary or multi-choice classification based on source image and textual query. |
| Approach: | They propose a task where a textual premise is the background presumption on each source image. |
| Outcome: | The proposed task is based on a dataset of 15,360 movie screenshots and human-curated premise templates from 6 pre-defined categories. |
Grounded PCFG Induction with Images (2020.aacl-main)
Copied to clipboard
| Challenge: | Recent work in unsupervised parsing has tried to incorporate visual information into learning, but results suggest that these models need linguistic bias to compete against models that only rely on text. |
| Approach: | They propose to use visual information from images for labeled parsing and compare them to existing models which only use text. |
| Outcome: | The proposed models achieve state-of-the-art results on multilingual induction datasets even without help from linguistic knowledge or pretrained image encoders. |
Probing Multimodal Embeddings for Linguistic Properties: the Visual-Semantic Case (2020.coling-main)
Copied to clipboard
| Challenge: | Semantic embeddings have advanced the state of the art for natural language processing tasks . but their inner workings are poorly understood and there is a shortage of analysis tools . |
| Approach: | They propose to extend visual-semantic embeddings to multimodal domains by defining probing tasks for embeddable image-caption pairs and testing them with classifiers. |
| Outcome: | The proposed probing tasks show up to 16% more accurate on visual-semantic embeddings compared to unimodal embedders . the proposed extensions to multimodal domains have been lauded as promising in natural language processing . |