Papers by Manishit Kundu
CaRVE: Critiquing and Refining Visual Elaborations for Figurative Language Illustrations (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing text-to-image frameworks for figurative illustration rely on proprietary models or human supervision to achieve adequate alignment. |
| Approach: | They propose a critique-driven framework that uses VLM feedback to refine visual elaborations for figurative image generation. |
| Outcome: | The proposed framework outperforms existing figurative image-to-text pipelines on human-supervised visual elaborations. |
Looking Beyond the Pixels: Evaluating Visual Metaphor Understanding in VLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Visual metaphors are a complex vision–language phenomenon that requires both perceptual and conceptual reasoning to understand. |
| Approach: | They introduce a visual metaphor dataset featuring 2177 synthetic and 350 human-annotated images and benchmark several SOTA VLMs on two tasks: Visual Metaphor Captioning (VMC) and Visual Metamorphosis VQA (VM-VQA). |
| Outcome: | The proposed model outperforms standard few-shot baselines on visual metaphors and VM-VQA tasks. |