Challenge: Recent research in language and vision has developed models for predicting and disambiguating verbs from images.
Approach: They propose a verb prediction model and visual sense disambiguation model for verbs . they ask whether the image regions a model identifies as salient correlate with human intuitions about visual verbs.
Outcome: The proposed model can predict verbs from images, but it is unclear to what extent it captures human intuitions about visual verbs.

Similar Papers

Representing Verbs with Visual Argument Vectors (2020.lrec-1)

Copied to clipboard

Challenge: Existing models for verb semantic similarities are based on linguistic data, but they do not register intuitive attributes.
Approach: They evaluated two textual distributional semantic models and a visual one to explore verb semantic similarities.
Outcome: The proposed models extract meaningful information and capture semantic similarity between verbs using visual distributional models.
Probing Image-Language Transformers for Verb Understanding (2021.findings-acl)

Copied to clipboard

Challenge: Multimodal image-language transformers have achieved impressive results on a variety of tasks that rely on fine-tuning.
Approach: They collect a dataset of image-sentence pairs consisting of 421 verbs . they evaluate pretrained image-language transformers and find they fail more in situations that require verb understanding compared to other parts of speech.
Outcome: The proposed model trains on a manually-annotated and smaller dataset does better on the task.
Visual Referring Expression Recognition: What Do Systems Actually Learn? (N18-2)

Copied to clipboard

Challenge: Existing systems for referring expression recognition ignore linguistic structure, instead relying on shallow correlations introduced by unintended biases in the data selection and annotation process.
Approach: They propose to use a system trained on the input image without the input referring expression to achieve a precision of 71.2% in top-2 predictions.
Outcome: The proposed model can achieve 71.2% accuracy on the input image without the input referring expression and 84.2% on the object category given the input.
A Psycholinguistic Evaluation of Language Models’ Sensitivity to Argument Roles (2024.findings-emnlp)

Copied to clipboard

Challenge: a systematic evaluation of large language models' sensitivity to argument roles is presented . a recent study shows that argument roles have a delayed impact on verb prediction in human sentence processing.
Approach: They propose to replicate psycholinguistic studies on human argument role processing . they find that language models are able to distinguish verbs that appear in plausible and implausible contexts .
Outcome: The proposed models are able to distinguish verbs that appear in plausible and implausible contexts, but none captures the same selective patterns that human comprehenders exhibit during real-time verb prediction.
Cross-lingual Visual Verb Sense Disambiguation (N19-1)

Copied to clipboard

Challenge: Recent work has shown that visual context improves cross-lingual sense disambiguation for nouns.
Approach: They extend their work to the task of cross-lingual verb sense disambiguation by using a dataset annotated with English, German, and Spanish verbs.
Outcome: The proposed model improves the results of a text-only machine translation system when used for a multimodal translation task.
Towards Explainable Evaluation of Language Models on the Semantic Similarity of Visual Concepts (2022.coling-1)

Copied to clipboard

Challenge: Recent advances in NLP research have focused on robustness and explainability issues of their evaluation strategies.
Approach: They propose to use pre-trained transformers to evaluate semantic similarity for visual vocabularies . they propose to provide explainable metrics for understanding the quality of retrieved instances .
Outcome: The proposed metrics highlight inabilities of widely used evaluation methods and highlight weaknesses in learned linguistic representations.
An Attentive Recurrent Model for Incremental Prediction of Sentence-final Verbs (2020.findings-emnlp)

Copied to clipboard

Challenge: Using an attention-based neural model, we can incrementally predict verbs on incomplete sentences in Japanese and German SOV sentences.
Approach: They propose a synonym-aware neural model to incrementally predict final verbs on incomplete sentences in Japanese and German SOV sentences.
Outcome: The proposed model outperforms existing models in predicting most frequent verbs in Japanese and German . larger datasets always help with predicting the sentencefinal verbs, suggesting larger dataset could be used to reduce translation latency.
Native Language Prediction from Gaze: a Reproducibility Study (2023.acl-srw)

Copied to clipboard

Challenge: Existing studies have shown that the linguistic properties of a speaker’s native language affect the cognitive processing of other languages.
Approach: They found that the correlation between eye movements and native language similarity may be more complex than the original study found.
Outcome: The proposed model shows that the correlation between eye movements and native language similarity may be more complex than the original study.
Describing Images Fast and Slow: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes (2024.eacl-long)

Copied to clipboard

Challenge: Existing models of visuo-linguistic variation are weak to moderately trained to capture such a variation in visual outputs.
Approach: They use a corpus of Dutch image descriptions with eye-tracking data to investigate the nature of the variation in visuo-linguistic signals.
Outcome: The proposed model lacks biases about what makes a stimulus complex for humans and what leads to variations in human outputs.
VROAV: Using Iconicity to Visually Represent Abstract Verbs (2020.lrec-1)

Copied to clipboard

Challenge: Visual languages like sign languages reveal enlightening patterns across signs of similar meanings, pointing towards the possibility of identifying clusters of iconic meanings.
Approach: a new verb classification system is proposed to visually represent 20 classes of abstract verbs.
Outcome: The proposed system could be used as a language learning aid or as linguistic comprehension tool for digital text.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations