Challenge: a long tradition of cognitive studies shows that the interplay between language and vision is complex.
Approach: They propose an approach to image description generation where visual processing is modelled sequentially.
Outcome: The proposed model exploits gaze-driven attention to produce better descriptions . it sheds light on human cognitive processes by comparing different ways of aligning gaze with language production.

Similar Papers

Cross-modal Coherence Modeling for Caption Generation (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for image captioning do not guarantee consistent image-text relations . current models do not provide enough data for training robust captioning models .
Approach: They use an annotation protocol specifically devised for capturing image–caption coherence relations to study image captioning.
Outcome: The proposed protocol improves image captioning models with coherence relations . the dataset is large enough to alleviate content hallucinations, the authors show .
Uncovering Limitations in Text-to-Image Generation: A Contrastive Approach with Structured Semantic Alignment (2023.findings-emnlp)

Copied to clipboard

Challenge: a new method for text-to-image generation models is proposed to address these limitations . SSA focuses on learning structured semantic embeddings across different modalities .
Approach: They propose a method to evaluate text-to-image generation models using structured semantic embeddings . they propose to learn mutated prompts by substituting words with equivalent or nonequivalent alternatives .
Outcome: The proposed method improves the measurement of semantic consistency of text-to-image generation models.
Seeing Eye-to-Eye: Cross-Modal Coherence Relations Inform Eye-gaze Patterns During Comprehension & Production (2024.lrec-main)

Copied to clipboard

Challenge: Xu and Stone et al., 2014, show eye movements are correlated with discourse goals but the relationship between eye movements and coherence is a missing link.
Approach: They propose an eye gaze pattern ranking algorithm and a semantic gaze visualization technique to study eye gaze patterns and coherence relations in multimodal language contexts.
Outcome: The proposed method combines eye-tracking and a semantic gaze visualization technique to study eye movements in multimodal language contexts.
An Annotation Approach for Social and Referential Gaze in Dialogue (2020.lrec-1)

Copied to clipboard

Challenge: Existing studies on eye gaze information focus on social functions and how it is used in reference resolution.
Approach: They propose an approach for annotating eye gaze considering its social and referential functions in multi-modal dialogue.
Outcome: The proposed annotation scheme is based on eye gaze behavior cues in human-human dialogues.
Controlling Reading Ease with Gaze-Guided Text Generation (2026.eacl-long)

Copied to clipboard

Challenge: Using a gaze-based model, we generate texts with controllable reading ease.
Approach: They propose a method that predicts gaze patterns to steer language model outputs towards eliciting certain reading behaviors by predicting eye-tracking measures.
Outcome: The proposed method generates texts with controllable reading ease using eye-tracking with native and non-native speakers of English.
Probing Cross-modal Semantics Alignment Capability from the Textual Perspective (2022.findings-emnlp)

Copied to clipboard

Challenge: In recent years, vision and language pre-training (VLP) models have advanced the state-of-the-art results in a variety of cross-modal downstream tasks.
Approach: They propose a new probing method that is based on image captioning to first empirically study the cross-modal semantics alignment of VLP models.
Outcome: The proposed method analyzes captions generated by five popular VLP models to reveal how well they align with visual words and how well these align with images.
ImageInWords: Unlocking Hyper-Detailed Image Descriptions (2024.emnlp-main)

Copied to clipboard

Challenge: generating accurate hyper-detailed image descriptions is challenging for vision-language models trained on web-scraped image-text.
Approach: They propose a data-centric framework for generating hyper-detailed image descriptions using web-scraped image-text.
Outcome: The proposed framework improves on human evaluations on the data, even with only 9k samples.
Cross-Lingual Representation Alignment Through Contrastive Image-Caption Tuning (2025.acl-short)

Copied to clipboard

Challenge: Multilingual alignment of sentence representations has mostly required bitexts to bridge the gap between languages.
Approach: They propose to use image captions to implicitly align text representations between languages to make them usable for cross-lingual Natural Language Understanding (NLU) and bitext retrieval.
Outcome: The proposed approach is usable for cross-lingual Natural Language Understanding (NLU) and bitext retrieval.
SNAG: Spoken Narratives and Gaze Dataset (P18-2)

Copied to clipboard

Challenge: Existing datasets that combine gaze and spoken descriptions of visual inputs are needed to provide insight into how humans process information and make decisions.
Approach: They propose a multimodal gaze and spoken descriptions dataset that can be used to label important image regions with appropriate linguistic labels.
Outcome: The proposed dataset can be used to label image regions with appropriate linguistic labels.
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies show that deep vision-only and language-only models project inputs into a partially aligned representational space.
Approach: They investigate whether a model's representational code is semantically shared . they find that alignment peaks in mid-to-late layers of both model types .
Outcome: a forced-choice "Pick-a-Pic" task shows human preferences for image-caption matches are mirrored in embedding spaces across vision-language model pairs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations