Papers by Christoffer Heckman
Aligning Images and Text with Semantic Role Labels for Fine-Grained Cross-Modal Understanding (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, image retrieval systems can retrieve relevant results for diverse inputs, but they do not provide a way to intentionally inject variety into the search results. |
| Approach: | They propose a multimodal dataset that combines semantic annotations with image bounding boxes. |
| Outcome: | The proposed system improves image retrieval performance and flexibility. |
CRAPES:Cross-modal Annotation Projection for Visual Semantic Role Labeling (2023.starsem-1)
Copied to clipboard
| Challenge: | Existing approaches to image comprehension limit the image to a single action, while text-based approaches label all actions in a sentence. |
| Approach: | They propose to expand GSR to follow more liberal text-based approach to action and participant identification. |
| Outcome: | The proposed approach improves image comprehension on a SWiG dataset by 28.6 points. |
ReCAP: Semantic Role Enhanced Caption Generation (2024.lrec-main)
Copied to clipboard
| Challenge: | Current vision language models lack specificity and overlook various aspects of the image. |
| Approach: | They propose to use semantic roles as control signals to guide captions to specific argument structures by focusing on specific objects and their associated semantic roles instead of general descriptions. |
| Outcome: | The proposed framework produces captions that exhibit enhanced quality, diversity, and controllability. |