Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human Gaze (2020.emnlp-main)
Copied to clipboard
| Challenge: | a long tradition of cognitive studies shows that the interplay between language and vision is complex. |
| Approach: | They propose an approach to image description generation where visual processing is modelled sequentially. |
| Outcome: | The proposed model exploits gaze-driven attention to produce better descriptions . it sheds light on human cognitive processes by comparing different ways of aligning gaze with language production. |
Similar Papers
Cross-modal Coherence Modeling for Caption Generation (2020.acl-main)
Copied to clipboard
| Challenge: | Existing methods for image captioning do not guarantee consistent image-text relations . current models do not provide enough data for training robust captioning models . |
| Approach: | They use an annotation protocol specifically devised for capturing image–caption coherence relations to study image captioning. |
| Outcome: | The proposed protocol improves image captioning models with coherence relations . the dataset is large enough to alleviate content hallucinations, the authors show . |
Uncovering Limitations in Text-to-Image Generation: A Contrastive Approach with Structured Semantic Alignment (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a new method for text-to-image generation models is proposed to address these limitations . SSA focuses on learning structured semantic embeddings across different modalities . |
| Approach: | They propose a method to evaluate text-to-image generation models using structured semantic embeddings . they propose to learn mutated prompts by substituting words with equivalent or nonequivalent alternatives . |
| Outcome: | The proposed method improves the measurement of semantic consistency of text-to-image generation models. |
Seeing Eye-to-Eye: Cross-Modal Coherence Relations Inform Eye-gaze Patterns During Comprehension & Production (2024.lrec-main)
Copied to clipboard
| Challenge: | Xu and Stone et al., 2014, show eye movements are correlated with discourse goals but the relationship between eye movements and coherence is a missing link. |
| Approach: | They propose an eye gaze pattern ranking algorithm and a semantic gaze visualization technique to study eye gaze patterns and coherence relations in multimodal language contexts. |
| Outcome: | The proposed method combines eye-tracking and a semantic gaze visualization technique to study eye movements in multimodal language contexts. |
An Annotation Approach for Social and Referential Gaze in Dialogue (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing studies on eye gaze information focus on social functions and how it is used in reference resolution. |
| Approach: | They propose an approach for annotating eye gaze considering its social and referential functions in multi-modal dialogue. |
| Outcome: | The proposed annotation scheme is based on eye gaze behavior cues in human-human dialogues. |
Controlling Reading Ease with Gaze-Guided Text Generation (2026.eacl-long)
Copied to clipboard
| Challenge: | Using a gaze-based model, we generate texts with controllable reading ease. |
| Approach: | They propose a method that predicts gaze patterns to steer language model outputs towards eliciting certain reading behaviors by predicting eye-tracking measures. |
| Outcome: | The proposed method generates texts with controllable reading ease using eye-tracking with native and non-native speakers of English. |
Probing Cross-modal Semantics Alignment Capability from the Textual Perspective (2022.findings-emnlp)
Copied to clipboard
| Challenge: | In recent years, vision and language pre-training (VLP) models have advanced the state-of-the-art results in a variety of cross-modal downstream tasks. |
| Approach: | They propose a new probing method that is based on image captioning to first empirically study the cross-modal semantics alignment of VLP models. |
| Outcome: | The proposed method analyzes captions generated by five popular VLP models to reveal how well they align with visual words and how well these align with images. |
ImageInWords: Unlocking Hyper-Detailed Image Descriptions (2024.emnlp-main)
Copied to clipboard
Roopal Garg, Andrea Burns, Burcu Karagol Ayan, Yonatan Bitton, Ceslee Montgomery, Yasumasa Onoe, Andrew Bunner, Ranjay Krishna, Jason Baldridge, Radu Soricut
| Challenge: | generating accurate hyper-detailed image descriptions is challenging for vision-language models trained on web-scraped image-text. |
| Approach: | They propose a data-centric framework for generating hyper-detailed image descriptions using web-scraped image-text. |
| Outcome: | The proposed framework improves on human evaluations on the data, even with only 9k samples. |
Cross-Lingual Representation Alignment Through Contrastive Image-Caption Tuning (2025.acl-short)
Copied to clipboard
| Challenge: | Multilingual alignment of sentence representations has mostly required bitexts to bridge the gap between languages. |
| Approach: | They propose to use image captions to implicitly align text representations between languages to make them usable for cross-lingual Natural Language Understanding (NLU) and bitext retrieval. |
| Outcome: | The proposed approach is usable for cross-lingual Natural Language Understanding (NLU) and bitext retrieval. |
SNAG: Spoken Narratives and Gaze Dataset (P18-2)
Copied to clipboard
| Challenge: | Existing datasets that combine gaze and spoken descriptions of visual inputs are needed to provide insight into how humans process information and make decisions. |
| Approach: | They propose a multimodal gaze and spoken descriptions dataset that can be used to label important image regions with appropriate linguistic labels. |
| Outcome: | The proposed dataset can be used to label image regions with appropriate linguistic labels. |
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies show that deep vision-only and language-only models project inputs into a partially aligned representational space. |
| Approach: | They investigate whether a model's representational code is semantically shared . they find that alignment peaks in mid-to-late layers of both model types . |
| Outcome: | a forced-choice "Pick-a-Pic" task shows human preferences for image-caption matches are mirrored in embedding spaces across vision-language model pairs. |