Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing evaluation methods for visual activity recognition systems fail to capture ambiguities in verb semantics and image interpretation. |
| Approach: | They propose a framework that constructs verb sense clusters to evaluate visual activity recognition systems. |
| Outcome: | The proposed framework provides a more robust evaluation of visual activity recognition systems. |
Similar Papers
Towards Explainable Evaluation of Language Models on the Semantic Similarity of Visual Concepts (2022.coling-1)
Copied to clipboard
Maria Lymperaiou, George Manoliadis, Orfeas Menis Mastromichalakis, Edmund G. Dervakos, Giorgos Stamou
| Challenge: | Recent advances in NLP research have focused on robustness and explainability issues of their evaluation strategies. |
| Approach: | They propose to use pre-trained transformers to evaluate semantic similarity for visual vocabularies . they propose to provide explainable metrics for understanding the quality of retrieved instances . |
| Outcome: | The proposed metrics highlight inabilities of widely used evaluation methods and highlight weaknesses in learned linguistic representations. |
Manual Clustering and Spatial Arrangement of Verbs for Multilingual Evaluation and Typology Analysis (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods to learn general language representations from large volumes of unlabeled text have been used to improve multilingual NLP. |
| Approach: | They propose to use a spatial arrangement method to generate large-scale evaluation datasets that balance cross-lingual alignment with language specificity. |
| Outcome: | The proposed method produces semantic verb classes and fine-grained similarity scores for nearly 130 thousand verb pairs. |
Spatial Multi-Arrangement for Clustering and Multi-way Similarity Dataset Construction (2020.lrec-1)
Copied to clipboard
Olga Majewska, Diana McCarthy, Jasper van den Bosch, Nikolaus Kriegeskorte, Ivan Vulić, Anna Korhonen
| Challenge: | Existing methods for creating large-scale semantic similarity resources are slow and expensive . a large verb similarity dataset is available for a number of verbs, but not for English. |
| Approach: | They propose a method for fast bottom-up creation of large-scale semantic similarity resources . they leverage semantic intuitions of native speakers and adapt a spatial multi-arrangement approach to lexical stimuli. |
| Outcome: | The proposed approach produces a large-scale verb similarity dataset containing similarity scores for 29,721 unique verb pairs and 825 target verbs. |
An Evaluation of Image-Based Verb Prediction Models against Human Eye-Tracking Data (N18-2)
Copied to clipboard
| Challenge: | Recent research in language and vision has developed models for predicting and disambiguating verbs from images. |
| Approach: | They propose a verb prediction model and visual sense disambiguation model for verbs . they ask whether the image regions a model identifies as salient correlate with human intuitions about visual verbs. |
| Outcome: | The proposed model can predict verbs from images, but it is unclear to what extent it captures human intuitions about visual verbs. |
Representing Verbs with Visual Argument Vectors (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing models for verb semantic similarities are based on linguistic data, but they do not register intuitive attributes. |
| Approach: | They evaluated two textual distributional semantic models and a visual one to explore verb semantic similarities. |
| Outcome: | The proposed models extract meaningful information and capture semantic similarity between verbs using visual distributional models. |
VROAV: Using Iconicity to Visually Represent Abstract Verbs (2020.lrec-1)
Copied to clipboard
| Challenge: | Visual languages like sign languages reveal enlightening patterns across signs of similar meanings, pointing towards the possibility of identifying clusters of iconic meanings. |
| Approach: | a new verb classification system is proposed to visually represent 20 classes of abstract verbs. |
| Outcome: | The proposed system could be used as a language learning aid or as linguistic comprehension tool for digital text. |
Acquiring Verb Classes Through Bottom-Up Semantic Verb Clustering (L18-1)
Copied to clipboard
| Challenge: | Existing methods for creating verbal classifications are limited or non-existent in most languages . a range of automatic verb classification approaches have been proposed, but high-quality resources are needed . |
| Approach: | They propose to use top-up semantic clustering to extract syntactic and semantic information from verbs in English, Polish and Croatian. |
| Outcome: | The proposed classifications in English, Polish and Croatian are compared with other languages. |
MM-R3: On (In-)Consistency of Vision-Language Models (VLMs) (2025.findings-acl)
Copied to clipboard
| Challenge: | a flurry of research has been conducted on the performance of state-of-the-art (SoTA) Vision Language Models (VLMs) on a variety of tasks. |
| Approach: | They propose a benchmarking tool to analyze performance of SoTA Vision Language Models (VLMs) on three tasks: Question Rephrasing, Image Restyling, and Context Reasoning. |
| Outcome: | The proposed model achieves absolute improvements of 5.7% and 12.5% on widely used VLMs such as BLIP-2 and LLaVa 1.5M in terms of consistency over their existing counterparts. |
Probing Image-Language Transformers for Verb Understanding (2021.findings-acl)
Copied to clipboard
| Challenge: | Multimodal image-language transformers have achieved impressive results on a variety of tasks that rely on fine-tuning. |
| Approach: | They collect a dataset of image-sentence pairs consisting of 421 verbs . they evaluate pretrained image-language transformers and find they fail more in situations that require verb understanding compared to other parts of speech. |
| Outcome: | The proposed model trains on a manually-annotated and smaller dataset does better on the task. |
Cross-lingual Visual Verb Sense Disambiguation (N19-1)
Copied to clipboard
| Challenge: | Recent work has shown that visual context improves cross-lingual sense disambiguation for nouns. |
| Approach: | They extend their work to the task of cross-lingual verb sense disambiguation by using a dataset annotated with English, German, and Spanish verbs. |
| Outcome: | The proposed model improves the results of a text-only machine translation system when used for a multimodal translation task. |