Papers by Claudio Greco

4 papers
Grounded Textual Entailment (C18-1)

Copied to clipboard

Challenge: Existing models for entailment analysis are not performing well in visual information-based models.
Approach: They propose to use a visual representation of the Textual Entailment task to compare visual-grounded models with a multimodal version of the SNLI dataset.
Outcome: The proposed model performs better when there is an image of the “world” or “situation” .
Psycholinguistics Meets Continual Learning: Measuring Catastrophic Forgetting in Visual Question Answering (P19-1)

Copied to clipboard

Challenge: Existing methods to overcome catastrophic forgetting in visual question answering models are inadequate, but have received little attention within natural language processing.
Approach: They devise a set of linguistically-informed visual question answering tasks motivated by psycholinguistics and investigate impact of task difficulty on continual learning.
Outcome: The proposed models differ in the types of questions they ask and show that task difficulty and order matter.
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding (2024.findings-emnlp)

Copied to clipboard

Challenge: Current Vision-Language Models (VLMs) focus on third-person view videos, neglecting the richness of egocentric perceptual experience.
Approach: They propose to use the Egocentric Video Understanding Dataset (EVUD) to train VLMs on video captioning and question answering tasks specific to egocentric videos.
Outcome: The proposed model outperforms open-source models including strong Socratic models using GPT-4 as a planner by 3.6% and outperformed Claude 3 and Gemini Pro Vision 1.0.
Be Different to Be Better! A Benchmark to Leverage the Complementarity of Language and Vision (2020.findings-emnlp)

Copied to clipboard

Challenge: BD2BB is a language and vision benchmark that requires multimodal models combine complementary information from the two modalities.
Approach: They propose a novel language and vision benchmark that requires multimodal models combine complementary information from both modalities.
Outcome: The proposed model is easy for humans, but poor for humans . it compares state-of-the-art models against human speakers to show that it performs well.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations