Papers by Claudio Greco
Grounded Textual Entailment (C18-1)
Copied to clipboard
Hoa Trong Vu, Claudio Greco, Aliia Erofeeva, Somayeh Jafaritazehjan, Guido Linders, Marc Tanti, Alberto Testoni, Raffaella Bernardi, Albert Gatt
| Challenge: | Existing models for entailment analysis are not performing well in visual information-based models. |
| Approach: | They propose to use a visual representation of the Textual Entailment task to compare visual-grounded models with a multimodal version of the SNLI dataset. |
| Outcome: | The proposed model performs better when there is an image of the “world” or “situation” . |
Psycholinguistics Meets Continual Learning: Measuring Catastrophic Forgetting in Visual Question Answering (P19-1)
Copied to clipboard
| Challenge: | Existing methods to overcome catastrophic forgetting in visual question answering models are inadequate, but have received little attention within natural language processing. |
| Approach: | They devise a set of linguistically-informed visual question answering tasks motivated by psycholinguistics and investigate impact of task difficulty on continual learning. |
| Outcome: | The proposed models differ in the types of questions they ask and show that task difficulty and order matter. |
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding (2024.findings-emnlp)
Copied to clipboard
Alessandro Suglia, Claudio Greco, Katie Baker, Jose Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, Oliver Lemon
| Challenge: | Current Vision-Language Models (VLMs) focus on third-person view videos, neglecting the richness of egocentric perceptual experience. |
| Approach: | They propose to use the Egocentric Video Understanding Dataset (EVUD) to train VLMs on video captioning and question answering tasks specific to egocentric videos. |
| Outcome: | The proposed model outperforms open-source models including strong Socratic models using GPT-4 as a planner by 3.6% and outperformed Claude 3 and Gemini Pro Vision 1.0. |
Be Different to Be Better! A Benchmark to Leverage the Complementarity of Language and Vision (2020.findings-emnlp)
Copied to clipboard
| Challenge: | BD2BB is a language and vision benchmark that requires multimodal models combine complementary information from the two modalities. |
| Approach: | They propose a novel language and vision benchmark that requires multimodal models combine complementary information from both modalities. |
| Outcome: | The proposed model is easy for humans, but poor for humans . it compares state-of-the-art models against human speakers to show that it performs well. |