An Evaluation of Image-Based Verb Prediction Models against Human Eye-Tracking Data (N18-2)
Copied to clipboard
| Challenge: | Recent research in language and vision has developed models for predicting and disambiguating verbs from images. |
| Approach: | They propose a verb prediction model and visual sense disambiguation model for verbs . they ask whether the image regions a model identifies as salient correlate with human intuitions about visual verbs. |
| Outcome: | The proposed model can predict verbs from images, but it is unclear to what extent it captures human intuitions about visual verbs. |
Similar Papers
Representing Verbs with Visual Argument Vectors (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing models for verb semantic similarities are based on linguistic data, but they do not register intuitive attributes. |
| Approach: | They evaluated two textual distributional semantic models and a visual one to explore verb semantic similarities. |
| Outcome: | The proposed models extract meaningful information and capture semantic similarity between verbs using visual distributional models. |
Probing Image-Language Transformers for Verb Understanding (2021.findings-acl)
Copied to clipboard
| Challenge: | Multimodal image-language transformers have achieved impressive results on a variety of tasks that rely on fine-tuning. |
| Approach: | They collect a dataset of image-sentence pairs consisting of 421 verbs . they evaluate pretrained image-language transformers and find they fail more in situations that require verb understanding compared to other parts of speech. |
| Outcome: | The proposed model trains on a manually-annotated and smaller dataset does better on the task. |
Visual Referring Expression Recognition: What Do Systems Actually Learn? (N18-2)
Copied to clipboard
| Challenge: | Existing systems for referring expression recognition ignore linguistic structure, instead relying on shallow correlations introduced by unintended biases in the data selection and annotation process. |
| Approach: | They propose to use a system trained on the input image without the input referring expression to achieve a precision of 71.2% in top-2 predictions. |
| Outcome: | The proposed model can achieve 71.2% accuracy on the input image without the input referring expression and 84.2% on the object category given the input. |
A Psycholinguistic Evaluation of Language Models’ Sensitivity to Argument Roles (2024.findings-emnlp)
Copied to clipboard
| Challenge: | a systematic evaluation of large language models' sensitivity to argument roles is presented . a recent study shows that argument roles have a delayed impact on verb prediction in human sentence processing. |
| Approach: | They propose to replicate psycholinguistic studies on human argument role processing . they find that language models are able to distinguish verbs that appear in plausible and implausible contexts . |
| Outcome: | The proposed models are able to distinguish verbs that appear in plausible and implausible contexts, but none captures the same selective patterns that human comprehenders exhibit during real-time verb prediction. |
Cross-lingual Visual Verb Sense Disambiguation (N19-1)
Copied to clipboard
| Challenge: | Recent work has shown that visual context improves cross-lingual sense disambiguation for nouns. |
| Approach: | They extend their work to the task of cross-lingual verb sense disambiguation by using a dataset annotated with English, German, and Spanish verbs. |
| Outcome: | The proposed model improves the results of a text-only machine translation system when used for a multimodal translation task. |
Towards Explainable Evaluation of Language Models on the Semantic Similarity of Visual Concepts (2022.coling-1)
Copied to clipboard
Maria Lymperaiou, George Manoliadis, Orfeas Menis Mastromichalakis, Edmund G. Dervakos, Giorgos Stamou
| Challenge: | Recent advances in NLP research have focused on robustness and explainability issues of their evaluation strategies. |
| Approach: | They propose to use pre-trained transformers to evaluate semantic similarity for visual vocabularies . they propose to provide explainable metrics for understanding the quality of retrieved instances . |
| Outcome: | The proposed metrics highlight inabilities of widely used evaluation methods and highlight weaknesses in learned linguistic representations. |
An Attentive Recurrent Model for Incremental Prediction of Sentence-final Verbs (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Using an attention-based neural model, we can incrementally predict verbs on incomplete sentences in Japanese and German SOV sentences. |
| Approach: | They propose a synonym-aware neural model to incrementally predict final verbs on incomplete sentences in Japanese and German SOV sentences. |
| Outcome: | The proposed model outperforms existing models in predicting most frequent verbs in Japanese and German . larger datasets always help with predicting the sentencefinal verbs, suggesting larger dataset could be used to reduce translation latency. |
Native Language Prediction from Gaze: a Reproducibility Study (2023.acl-srw)
Copied to clipboard
| Challenge: | Existing studies have shown that the linguistic properties of a speaker’s native language affect the cognitive processing of other languages. |
| Approach: | They found that the correlation between eye movements and native language similarity may be more complex than the original study found. |
| Outcome: | The proposed model shows that the correlation between eye movements and native language similarity may be more complex than the original study. |
Describing Images Fast and Slow: Quantifying and Predicting the Variation in Human Signals during Visuo-Linguistic Processes (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing models of visuo-linguistic variation are weak to moderately trained to capture such a variation in visual outputs. |
| Approach: | They use a corpus of Dutch image descriptions with eye-tracking data to investigate the nature of the variation in visuo-linguistic signals. |
| Outcome: | The proposed model lacks biases about what makes a stimulus complex for humans and what leads to variations in human outputs. |
VROAV: Using Iconicity to Visually Represent Abstract Verbs (2020.lrec-1)
Copied to clipboard
| Challenge: | Visual languages like sign languages reveal enlightening patterns across signs of similar meanings, pointing towards the possibility of identifying clusters of iconic meanings. |
| Approach: | a new verb classification system is proposed to visually represent 20 classes of abstract verbs. |
| Outcome: | The proposed system could be used as a language learning aid or as linguistic comprehension tool for digital text. |