| Challenge: | a dataset of 41k sentences describes fine-grained differences between photographs of birds . human observers are adept at making fine-grain comparisons, but sometimes require aid in distinguishing visually similar classes. |
| Approach: | They propose a model that generates comparative language from a dataset of 41k sentences describing fine-grained differences between photographs of birds. |
| Outcome: | The proposed model can explain differences in visual embedding space using natural language . it evaluates the results with humans who must use the descriptions to distinguish real images . |
Similar Papers
A Corpus for Reasoning about Natural Language Grounded in Photographs (P19-1)
Copied to clipboard
| Challenge: | a dataset for visual reasoning with natural language and images is available. |
| Approach: | They propose a dataset for joint reasoning about natural language and images . they crowdsource 107,292 examples of English sentences paired with web photographs . |
| Outcome: | The proposed dataset combines 107,292 examples of English sentences with web photographs . Qualitative analysis shows the data requires compositional joint reasoning . |
African or European Swallow? Benchmarking Large Vision-Language Models for Fine-Grained Object Classification (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent Large Vision Language Models demonstrate impressive abilities on image understanding and reasoning tasks. |
| Approach: | They propose a benchmark for fine-grained object classification that is difficult to evaluate . they benchmark 12 public LVLMs on and show CLIP models exhibit better performance . |
| Outcome: | The proposed model improves on 12 public LVLMs on image understanding and reasoning tasks. |
Learning to Describe Differences Between Pairs of Similar Images (D18-1)
Copied to clipboard
| Challenge: | Using crowd-sourced annotations, we generate text descriptions of differences between two images . we use a dataset to generate concise and fluent descriptions of visual data . |
| Approach: | They propose a task of automatically generating text to describe the differences between two images . they crowd-sourced the difference descriptions for pairs of images extracted from video-surveillance footage . |
| Outcome: | The proposed model outperforms models that use attention alone for single-sentence generation and multi-sentent generation. |
MISMATCH: Fine-grained Evaluation of Machine-generated Text with Mismatch Error Types (2023.findings-acl)
Copied to clipboard
Keerthiram Murugesan, Sarathkrishna Swaminathan, Soham Dan, Subhajit Chaudhury, Chulaka Gunasekara, Maxwell Crouse, Diwakar Mahajan, Ibrahim Abdelaziz, Achille Fokoue, Pavan Kapanipathi, Salim Roukos, Alexander Gray
| Challenge: | Existing evaluation metrics for machine text are inadequate to capture quality of text . a recent study has focused on task-specific evaluation metrics or on properties of machine-generated text based on mismatch errors . |
| Approach: | They propose a new evaluation scheme based on fine-grained mismatch errors . they propose 13 mismatch error types to guide the model for better prediction of human judgments . |
| Outcome: | The proposed evaluation scheme is based on mismatch errors in 7 NLP tasks . the mismatch error types guide the model for better prediction of human judgments . |
Fine-grained Medical Vision-Language Representation Learning for Radiology Report Generation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to learn medical vision-language representations by contrasting images with entire reports are not effective. |
| Approach: | They propose a phenotype-driven medical vision-language representation learning framework to bridge the gap between visual and textual modalities for improved text-oriented generation. |
| Outcome: | The proposed framework bridges the gap between visual and textual modalities for improved radiology report generation. |
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies show that deep vision-only and language-only models project inputs into a partially aligned representational space. |
| Approach: | They investigate whether a model's representational code is semantically shared . they find that alignment peaks in mid-to-late layers of both model types . |
| Outcome: | a forced-choice "Pick-a-Pic" task shows human preferences for image-caption matches are mirrored in embedding spaces across vision-language model pairs. |
How Pre-trained Word Representations Capture Commonsense Physical Comparisons (D19-60)
Copied to clipboard
| Challenge: | Pre-trained word representations capture common sense on physical properties such as size and weight. |
| Approach: | They investigate whether pre-trained representations capture comparisons and find they have higher accuracy than previous approaches. |
| Outcome: | The proposed models learn a consistent ordering over all the objects in the comparisons. |
Face2Text: Collecting an Annotated Image Description Corpus for the Generation of Rich Face Descriptions (L18-1)
Copied to clipboard
Albert Gatt, Marc Tanti, Adrian Muscat, Patrizia Paggio, Reuben A Farrugia, Claudia Borg, Kenneth P Camilleri, Michael Rosner, Lonneke van der Plas
| Challenge: | a crowdsourcing study has been conducted to generate rich textual descriptions of human faces . the aim is to investigate how users describe images of human face images . |
| Approach: | They propose to extend the problem of automatically generating text from images to face description . they conducted an annotation study on a subset of the corpus to gain a better understanding of the variation they find in face descriptions . |
| Outcome: | The proposed corpus is based on images taken in the wild and is expected to be large enough to support non-trivial machine learning work on the automated description of faces. |
Visual Grounding Helps Learn Word Meanings in Low-Data Regimes (2024.naacl-long)
Copied to clipboard
| Challenge: | Modern neural language models (LMs) require distinctly un-human-like ways to achieve these results. |
| Approach: | They train a diverse set of LM architectures with and without auxiliary visual supervision on datasets of varying scales. |
| Outcome: | The proposed models exhibit better learning of syntactic categories, lexical relations, semantic features, word similarity and alignment with human neural representations. |
‘Lighter’ Can Still Be Dark: Modeling Comparative Color Descriptions (P18-2)
Copied to clipboard
| Challenge: | Multimodal approaches to object recognition ground adjectives and nouns from text using comparative adjectives. |
| Approach: | They propose a new paradigm of grounding comparative adjectives within the realm of color descriptions by using a vector model. |
| Outcome: | The proposed model generates representations of comparative adjectives with an average accuracy of 0.65 cosine similarity to the desired direction of change. |