Challenge: a multilingual study examines how vision constrains linguistic choice . we use existing annotations to investigate the effect of different visual conditions on numeral expressions in captions .
Approach: They propose a method that leverages existing corpora of images with captions written by native speakers to constrain linguistic choice.
Outcome: The proposed method covers four languages and five linguistic properties, including verb transitivity and use of numerals.

Similar Papers

Learning Translations via Images with a Massively Multilingual Image Dataset (P18-1)

Copied to clipboard

Challenge: Existing datasets for learning translations of words are limited to a few high-resource languages and unrealistically easy settings.
Approach: They propose a large-scale multilingual corpus of images labeled with the word they represent to facilitate translation research.
Outcome: The proposed method improves on an unsupervised technique that has been limited to a few languages and unrealistic settings.
MultiSubs: A Large-scale Multimodal and Multilingual Dataset (2022.lrec-1)

Copied to clipboard

Challenge: a large-scale multimodal and multilingual dataset is used to facilitate research on visual grounding of words to images in their contextual usage in language.
Approach: They propose a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language.
Outcome: The proposed dataset will facilitate research on visual grounding of words in their contextual usage in language.
Cross-Lingual and Cross-Cultural Variation in Image Descriptions (2025.naacl-long)

Copied to clipboard

Challenge: Behavioural and cognitive studies report cultural effects on perception, but these are limited in scope and hard to replicate.
Approach: They develop a method to accurately identify entities mentioned in captions and present in images, then measure how they vary across languages.
Outcome: The proposed method corroborates previous studies showing that languages that are geographically or genetically closer mention entities more frequently than others.
Naming, Describing, and Quantifying Visual Objects in Humans and LLMs (2024.acl-short)

Copied to clipboard

Challenge: Recent work has highlighted that speakers display a wide range of variability when asked to utter sentences, resulting in inter-speaker variability but also variability over time for the same speaker.
Approach: They evaluate Vision & Language Large Language Models (VLLMs) on three categories where humans show great subjective variability concerning the distribution over plausible labels.
Outcome: The proposed models can mimic human distributions over plausible labels, but fail to assign quantifiers, a task that requires more accurate, high-level reasoning.
Images in Language Space: Exploring the Suitability of Large Language Models for Vision & Language Tasks (2023.findings-acl)

Copied to clipboard

Challenge: Large language models have demonstrated robust performance on various language tasks using zero-shot or few-shot learning paradigms.
Approach: They propose to use open-source, open-access language models to make visual input accessible to the model using separate verbalisation models.
Outcome: The proposed model can handle visual input but also require strong reasoning component.
Multilingual Image Corpus – Towards a Multimodal and Multilingual Dataset (2022.lrec-1)

Copied to clipboard

Challenge: The goal of the project Multilingual Image Corpus is to provide a large image dataset with annotated objects and object descriptions in 24 languages.
Approach: They propose to provide a large image dataset with annotated objects and object descriptions in 24 languages.
Outcome: The project provides a large image dataset with annotated objects and object descriptions in 24 languages.
Representing Verbs with Visual Argument Vectors (2020.lrec-1)

Copied to clipboard

Challenge: Existing models for verb semantic similarities are based on linguistic data, but they do not register intuitive attributes.
Approach: They evaluated two textual distributional semantic models and a visual one to explore verb semantic similarities.
Outcome: The proposed models extract meaningful information and capture semantic similarity between verbs using visual distributional models.
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies show that deep vision-only and language-only models project inputs into a partially aligned representational space.
Approach: They investigate whether a model's representational code is semantically shared . they find that alignment peaks in mid-to-late layers of both model types .
Outcome: a forced-choice "Pick-a-Pic" task shows human preferences for image-caption matches are mirrored in embedding spaces across vision-language model pairs.
Scaling in Cognitive Modelling: a Multilingual Approach to Human Reading Times (2023.acl-short)

Copied to clipboard

Challenge: Neural language models provide conditional probability distributions over the lexicon that are predictive of human processing times.
Approach: They propose to use a transformer-based model to generate probabilistic estimates that are less predictive of early eye-tracking measurements reflecting lexical access and early semantic integration.
Outcome: The proposed models show that larger models capture late eye-tracking measurements that reflect the full integration of a word into the current language context.
Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Prior studies show that large language models map multilingual content into English-aligned representations at intermediate layers before projecting them back into target-language token spaces in the later layers.
Approach: They propose a method to identify and manipulate dimensions that are sparse and sparsity-based . they propose to use as few as 50 sentences of either parallel or monolingual data to manipulate these dimensions .
Outcome: Experiments on a multilingual generation control task show the interpretability of these dimensions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations