| Challenge: | Current studies on image captioning focus on single image, but there are no effective models for generating relational captions for two images. |
| Approach: | They propose a language-guided image editing dataset that contains real image pairs with corresponding editing instructions. |
| Outcome: | The proposed model outperforms baseline and existing methods on two datasets. |
Similar Papers
Improving Image Captioning with Better Use of Caption (2020.acl-main)
Copied to clipboard
| Challenge: | Existing approaches to image captioning focus on visual attention, but many do not. |
| Approach: | They propose a framework that explores semantics available in captions and leverages that to enhance both image representation and caption generation. |
| Outcome: | The proposed framework outperforms baselines on the MSCOCO dataset and is state-of-the-art under a wide range of evaluation metrics. |
Cross-Lingual Representation Alignment Through Contrastive Image-Caption Tuning (2025.acl-short)
Copied to clipboard
| Challenge: | Multilingual alignment of sentence representations has mostly required bitexts to bridge the gap between languages. |
| Approach: | They propose to use image captions to implicitly align text representations between languages to make them usable for cross-lingual Natural Language Understanding (NLU) and bitext retrieval. |
| Outcome: | The proposed approach is usable for cross-lingual Natural Language Understanding (NLU) and bitext retrieval. |
Learning to Relate from Captions and Bounding Boxes (P19-1)
Copied to clipboard
| Challenge: | Existing methods for classifying images without supervision are limited. |
| Approach: | They propose a top-down attention mechanism to align entities in captions to objects in the image and leverage the syntactic structure of captions for alignment. |
| Outcome: | The proposed model achieves a recall@50 of 15% and recall@100 of 25% on the relationships present in the image and predicts relations that are not present in captions. |
Building Joint Relationship Attention Network for Image-Text Generation (2022.coling-1)
Copied to clipboard
| Challenge: | et al., 2017) focus on visual features individually, while ignoring relationship information among image features that provides important guidance for generating sentences. |
| Approach: | They propose a joint relationship attention network that explores the relationships among image features. |
| Outcome: | The proposed method achieves state-of-the-art performance on large-scale datasets and on Flickr30k datasets. |
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies show that deep vision-only and language-only models project inputs into a partially aligned representational space. |
| Approach: | They investigate whether a model's representational code is semantically shared . they find that alignment peaks in mid-to-late layers of both model types . |
| Outcome: | a forced-choice "Pick-a-Pic" task shows human preferences for image-caption matches are mirrored in embedding spaces across vision-language model pairs. |
Visually-Enhanced Phrase Understanding (2023.findings-acl)
Copied to clipboard
| Challenge: | Large-scale vision-language pre-training models generate high-quality textual representations, which often outperform models that are purely text-based, such as BERT. |
| Approach: | They propose to utilize both textual and visual encoders of multi-modal pre-trained models to enhance language understanding tasks by generating an image associated with a textual prompt. |
| Outcome: | The proposed method outperforms models that are purely text-based on visual and textual understanding tasks and significantly improves the entity clustering task. |
StereoRel: Relational Triple Extraction from a Stereoscopic Perspective (2021.acl-long)
Copied to clipboard
| Challenge: | Existing methods for relational triple extraction still face challenges, including information loss and error propagation. |
| Approach: | They propose a model which maps relational triples to a three-dimensional space and leverages three decoders to extract them. |
| Outcome: | The proposed model outperforms the baselines on five public datasets. |
Modeling Complex Semantics Relation with Contrastively Fine-Tuned Relational Encoders (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for learning relational embeddings fail to capture nuanced representations and rich semantics. |
| Approach: | They propose different relational encoders designed to capture diverse relational aspects and semantic properties of entity pairs. |
| Outcome: | The proposed encoders capture diverse relational aspects and semantic properties of entity pairs. |
Image Description Dataset for Language Learners (2022.lrec-1)
Copied to clipboard
| Challenge: | Language learners are limited by the number of texts or speech they are asked to answer . automatic assessment of image descriptions requires a system that depends on both the learner's native language and the target language. |
| Approach: | They propose a dataset that consists of images, their descriptions, and assessment annotations . they propose 'automatic error correction' task that encodes multimodal information from a learner sentence with an image and accurately decodes a corrected sentence. |
| Outcome: | The proposed model can revise errors that cannot be revised without an image. |
Visually Grounded Reasoning across Languages and Cultures (2021.emnlp-main)
Copied to clipboard
| Challenge: | a new protocol allows for a multilingual hierarchy of concepts and images based on native speakers . the results suggest that the current models are not robust enough to handle multilingual data . |
| Approach: | They propose a protocol to construct an ImageNet-style hierarchy representative of more languages and cultures. |
| Outcome: | The proposed protocol lets the selection of concepts and images be entirely driven by native speakers, rather than scraping them automatically. |