Multimodal Logical Inference System for Visual-Textual Entailment (P19-2)

Copied to clipboard

Challenge: Recent studies of multimodal inference provide challenging tasks such as visual question answering and visual reasoning.
Approach: They propose an unsupervised multimodal logical inference system that can prove entailment relations between texts and images by combing semantic parsing and theorem proving.
Outcome: The proposed system can handle semantically complex sentences for visual-textual inference.

Similar Papers

Recognizing Multimodal Entailment (2021.acl-tutorials)

Copied to clipboard

Challenge: This tutorial introduces the multimodal entailment task for detecting semantic alignments . the task requires fine-grained understanding of visual and linguistic semantics questions .
Approach: This tutorial introduces the multimodal entailment task to machine learning . it introduces a dataset for recognizing multimodal alignments .
Outcome: This tutorial introduces the multimodal entailment task . it can be useful for detecting semantic alignments when a single modality alone is not enough .
Grounded Textual Entailment (C18-1)

Copied to clipboard

Challenge: Existing models for entailment analysis are not performing well in visual information-based models.
Approach: They propose to use a visual representation of the Textual Entailment task to compare visual-grounded models with a multimodal version of the SNLI dataset.
Outcome: The proposed model performs better when there is an image of the “world” or “situation” .
Visual-Textual Entailment with Quantities Using Model Checking and Knowledge Injection (2024.lrec-main)

Copied to clipboard

Challenge: Visual-textual entailment (VTE) is a critical task in multimodal inference.
Approach: They propose a visual-textual entailment system that solves VTE tasks with quantities and negation.
Outcome: The proposed system solves visual-textual entailment tasks with quantities and negation more robustly than previous approaches.
Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding (2026.findings-eacl)

Copied to clipboard

Challenge: figurative language is essential for expressing intent, emotion, and perspective . figural language is often dependent on Styles Reasoning, causing incongruities between expressions .
Approach: They propose a framework that induces reasoning capabilities to compact vision–language models . figurative language is essential in expressing intent, emotion, and perspective .
Outcome: The proposed framework can interpret multimodal figurative language, provide transparent reasoning traces, and generalize across multiple figurativ styles.
Premise-based Multimodal Reasoning: Conditional Inference on Joint Textual and Visual Clues (2022.acl-long)

Copied to clipboard

Challenge: Existing work in vision language cross-modal reasoning uses binary or multi-choice classification based on source image and textual query.
Approach: They propose a task where a textual premise is the background presumption on each source image.
Outcome: The proposed task is based on a dataset of 15,360 movie screenshots and human-curated premise templates from 6 pre-defined categories.
Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you! (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models lack this active understanding capacity, limiting their applicability in real-world scenarios.
Approach: They propose a benchmark to assess the impact of multimodal inputs on lexical ambiguities.
Outcome: The proposed benchmark assesses the impact of multimodal inputs on lexical ambiguities.
Probing Logical Reasoning of MLLMs in Scientific Diagrams (2025.emnlp-main)

Copied to clipboard

Challenge: logical reasoning is key to real-world applications like science education, environmental monitoring, and medical diagnostics.
Approach: They construct visual questions that follow seven structured templates with progressively more complex reasoning involved.
Outcome: The proposed models perform logical inferences based on visual information.
A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Multimodal mathematical Reasoning (MMR) has attracted increasing attention for its ability to solve mathematical problems involving both textual and visual modalities.
Approach: They review the theoretical frameworks of multimodal reasoning and examine the challenges they face in visual math tasks.
Outcome: The proposed models can solve problems involving both textual and visual modalities.
A Probabilistic Model for Joint Learning of Word Embeddings from Texts and Images (D18-1)

Copied to clipboard

Challenge: Existing approaches combine language and perception to infer word embeddings . however, the embeddables produced by such models do not reflect the actual word representations.
Approach: They propose a probabilistic model that integrates linguistic and perceptual inputs to explain observed word-context pairs in a text corpus.
Outcome: The proposed model achieves competitive or stronger results on tasks of assessing pairwise word similarity and image/caption retrieval compared to other state-of-the-art models.
Multimodal Causal Reasoning Benchmark: Challenging Multimodal Large Language Models to Discern Causal Links Across Modalities (2025.findings-acl)

Copied to clipboard

Challenge: Existing MLLMs lack robustness in multimodal causal reasoning compared to their performance in textual settings.
Approach: They propose a novel multimodal chain-of-thought (CoT) reasoning benchmark that leverages siamese images and text pairs to challenge MLLMs.
Outcome: The proposed benchmark leverages siamese images and text pairs to challenge MLLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations