Papers with Flickr30K

7 papers
Learning to Represent Image and Text with Denotation Graph (2020.emnlp-main)

Copied to clipboard

Challenge: Recent advances in learning representations of visual and language information have been a problem with many applications.
Approach: They propose to extract visual expressions from images aligned with linguistic expressions that describe the images to learn representations from implicit expressions.
Outcome: The proposed representations lead to stronger empirical results on downstream tasks of cross-modal image retrieval, referring expression, and compositional attribute-object recognition.
The Role of Data Curation in Image Captioning (2024.eacl-long)

Copied to clipboard

Challenge: Existing image captioning models treat all samples equally, neglecting mismatched data . Several other techniques have relied on curriculum learning strategies to adapt learning to the difficulty of the task.
Approach: They propose to actively curate difficult samples in datasets using curriculum learning strategies to improve captioning models.
Outcome: The proposed methods outperform existing models on the Flickr30K and COCO datasets.
Improving Visual-Semantic Embedding with Adaptive Pooling and Optimization Objective (2023.eacl-main)

Copied to clipboard

Challenge: Recent VSE models combine simple pooling methods with hard triplet loss to improve performance.
Approach: They propose an adaptive pooling strategy that allows the model to learn how to aggregate features through a combination of simple pooling methods.
Outcome: The proposed strategy outperforms current state-of-the-art systems on image-to-text and text-toimage retrieval.
Leaner and Faster: Two-Stage Model Compression for Lightweight Text-Image Retrieval (2022.naacl-main)

Copied to clipboard

Challenge: Existing text-image approaches use pre-trained vision-language representations for text retrieval . however, these models pose non-trivial memory requirements and substantial indexing time .
Approach: They propose a framework to compress large pre-trained dual-encoders for lightweight text-image retrieval.
Outcome: The proposed model performs better on Flickr30K and MSCOCO benchmarks than the original full model on mobile devices.
An Anchor-based Relative Position Embedding Method for Cross-Modal Tasks (2022.emnlp-main)

Copied to clipboard

Challenge: Position Embedding (PE) is essential for transformer to capture the sequence ordering of input tokens.
Approach: They propose a unified position embedding method that bridges the semantic gap between modalities and embeds the anchor-based distance to guide computation of cross-attention.
Outcome: The proposed method obtains new SOTA results on a wide range of benchmarks.
Generative Text-to-Image Retrieval via Hierarchical Identifiers and Semantic Internalization (2026.findings-acl)

Copied to clipboard

Challenge: Existing text-to-image retrieval methods suffer from limited semantic discriminability, alignment bias, and closed-set restrictions.
Approach: They propose a framework for semantic internalization for Generative Multimodal Alignment . they construct multi-granularity hierarchical identifiers to ensure unique, semantically consistent image representations .
Outcome: The proposed framework outperforms state-of-the-art frameworks on Flickr30K and MS-COCO datasets . it achieves average Recall@1, Recall @5, and Recall_10 improvements of 10.65%, 8.50%, and 7.00% .
Visual-Linguistic Dependency Encoding for Image-Text Retrieval (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to image-text retrieval ignore semantic discrepancies caused by syntactic structure in natural language expressions and relationships among visual entities.
Approach: They propose a visual-linguistic dependency encoder framework which explicitly models the dependency information among textual words and interaction patterns between image regions.
Outcome: The proposed framework outperforms existing methods on a vision-linguistic compositional structure reasoning dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations