Papers with L&V
WikiScenes with Descriptions: Aligning Paragraphs and Sentences with Images in Wikipedia Articles (2024.starsem-1)
Copied to clipboard
| Challenge: | Existing work on processing image-text alignment in multimodal documents has been unsupervised, facing the challenge of missing evaluation and training data. |
| Approach: | They propose to provide one of the first datasets that provides ground-truth annotations of image-text alignments in multi-paragraph multi-image articles. |
| Outcome: | The proposed dataset can be used to study phenomena of visual language grounding in longer documents and assess retrieval capabilities of language models trained on captioning data. |
Know What You Don’t Know: Modeling a Pragmatic Speaker that Refers to Objects of Unknown Categories (P19-1)
Copied to clipboard
| Challenge: | a lot of recent and traditional research on pragmatically informative object descriptions has focused on the task of correctly labelling objects of novel categories. |
| Approach: | They extend a neural generator to become a pragmatic speaker reasoning about uncertain object categories. |
| Outcome: | The proposed model improves the accuracy of the listener's communication with unfamiliar objects. |
Humans Meet Models on Object Naming: A New Dataset and Analysis (2020.coling-main)
Copied to clipboard
| Challenge: | Existing object naming datasets that use only images with a bounding box are noisy . a human-like model behavior is not stable across domains, a study finds . |
| Approach: | They use MN v2 to verify object naming datasets with dozens of valid names per object . they find that human-like model behavior is not stable across domains . |
| Outcome: | The proposed model confuses people and clothing objects more frequently than humans do. |