Neural Naturalist: Generating Fine-Grained Image Comparisons (D19-1)

Copied to clipboard

Challenge: a dataset of 41k sentences describes fine-grained differences between photographs of birds . human observers are adept at making fine-grain comparisons, but sometimes require aid in distinguishing visually similar classes.
Approach: They propose a model that generates comparative language from a dataset of 41k sentences describing fine-grained differences between photographs of birds.
Outcome: The proposed model can explain differences in visual embedding space using natural language . it evaluates the results with humans who must use the descriptions to distinguish real images .

Similar Papers

A Corpus for Reasoning about Natural Language Grounded in Photographs (P19-1)

Copied to clipboard

Challenge: a dataset for visual reasoning with natural language and images is available.
Approach: They propose a dataset for joint reasoning about natural language and images . they crowdsource 107,292 examples of English sentences paired with web photographs .
Outcome: The proposed dataset combines 107,292 examples of English sentences with web photographs . Qualitative analysis shows the data requires compositional joint reasoning .
African or European Swallow? Benchmarking Large Vision-Language Models for Fine-Grained Object Classification (2024.emnlp-main)

Copied to clipboard

Challenge: Recent Large Vision Language Models demonstrate impressive abilities on image understanding and reasoning tasks.
Approach: They propose a benchmark for fine-grained object classification that is difficult to evaluate . they benchmark 12 public LVLMs on and show CLIP models exhibit better performance .
Outcome: The proposed model improves on 12 public LVLMs on image understanding and reasoning tasks.
Learning to Describe Differences Between Pairs of Similar Images (D18-1)

Copied to clipboard

Challenge: Using crowd-sourced annotations, we generate text descriptions of differences between two images . we use a dataset to generate concise and fluent descriptions of visual data .
Approach: They propose a task of automatically generating text to describe the differences between two images . they crowd-sourced the difference descriptions for pairs of images extracted from video-surveillance footage .
Outcome: The proposed model outperforms models that use attention alone for single-sentence generation and multi-sentent generation.
MISMATCH: Fine-grained Evaluation of Machine-generated Text with Mismatch Error Types (2023.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics for machine text are inadequate to capture quality of text . a recent study has focused on task-specific evaluation metrics or on properties of machine-generated text based on mismatch errors .
Approach: They propose a new evaluation scheme based on fine-grained mismatch errors . they propose 13 mismatch error types to guide the model for better prediction of human judgments .
Outcome: The proposed evaluation scheme is based on mismatch errors in 7 NLP tasks . the mismatch error types guide the model for better prediction of human judgments .
Fine-grained Medical Vision-Language Representation Learning for Radiology Report Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods to learn medical vision-language representations by contrasting images with entire reports are not effective.
Approach: They propose a phenotype-driven medical vision-language representation learning framework to bridge the gap between visual and textual modalities for improved text-oriented generation.
Outcome: The proposed framework bridges the gap between visual and textual modalities for improved radiology report generation.
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies show that deep vision-only and language-only models project inputs into a partially aligned representational space.
Approach: They investigate whether a model's representational code is semantically shared . they find that alignment peaks in mid-to-late layers of both model types .
Outcome: a forced-choice "Pick-a-Pic" task shows human preferences for image-caption matches are mirrored in embedding spaces across vision-language model pairs.
How Pre-trained Word Representations Capture Commonsense Physical Comparisons (D19-60)

Copied to clipboard

Challenge: Pre-trained word representations capture common sense on physical properties such as size and weight.
Approach: They investigate whether pre-trained representations capture comparisons and find they have higher accuracy than previous approaches.
Outcome: The proposed models learn a consistent ordering over all the objects in the comparisons.
Face2Text: Collecting an Annotated Image Description Corpus for the Generation of Rich Face Descriptions (L18-1)

Copied to clipboard

Challenge: a crowdsourcing study has been conducted to generate rich textual descriptions of human faces . the aim is to investigate how users describe images of human face images .
Approach: They propose to extend the problem of automatically generating text from images to face description . they conducted an annotation study on a subset of the corpus to gain a better understanding of the variation they find in face descriptions .
Outcome: The proposed corpus is based on images taken in the wild and is expected to be large enough to support non-trivial machine learning work on the automated description of faces.
Visual Grounding Helps Learn Word Meanings in Low-Data Regimes (2024.naacl-long)

Copied to clipboard

Challenge: Modern neural language models (LMs) require distinctly un-human-like ways to achieve these results.
Approach: They train a diverse set of LM architectures with and without auxiliary visual supervision on datasets of varying scales.
Outcome: The proposed models exhibit better learning of syntactic categories, lexical relations, semantic features, word similarity and alignment with human neural representations.
‘Lighter’ Can Still Be Dark: Modeling Comparative Color Descriptions (P18-2)

Copied to clipboard

Challenge: Multimodal approaches to object recognition ground adjectives and nouns from text using comparative adjectives.
Approach: They propose a new paradigm of grounding comparative adjectives within the realm of color descriptions by using a vector model.
Outcome: The proposed model generates representations of comparative adjectives with an average accuracy of 0.65 cosine similarity to the desired direction of change.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations