Papers by Somdeb Sarkhel
Question Modifiers in Visual Question Answering (2022.lrec-1)
Copied to clipboard
| Challenge: | Visual Question Answering (VQA) is a multi-disciplinary task that requires integration of several key disciplines. |
| Approach: | They develop a model that adds modifiers to questions based on object properties and spatial relationships using Amazon Mechanical Turk data. |
| Outcome: | The proposed model can improve when questions are modified to include more details. |
TAME-RD: Text Assisted Replication of Image Multi-Adjustments for Reverse Designing (2024.findings-acl)
Copied to clipboard
Pooja Guhan, Uttaran Bhattacharya, Somdeb Sarkhel, Vahid Azizi, Xiang Chen, Saayan Mitra, Aniket Bera, Dinesh Manocha
| Challenge: | a new model to reverse design images can be used to replicate image edits on other images based on human instructions in natural language . a study of a dataset of 100K source and edited images shows improvements in accuracy and concordance correlation coefficient . |
| Approach: | They propose a reverse-designing model that automatically learns from image editing operations and natural language instructions to learn fully specified edit operations. |
| Outcome: | The proposed model improves accuracy and concordance correlation scores on multiple datasets. |
VISIAR: Empower MLLM for Visual Story Ideation (2025.findings-acl)
Copied to clipboard
Zhaoyang Xia, Somdeb Sarkhel, Mehrab Tanjim, Stefano Petrangeli, Ishita Dasgupta, Yuxiao Chen, Jinxuan Xu, Di Liu, Saayan Mitra, Dimitris N. Metaxas
| Challenge: | Existing literature on visual storytelling has not explored the ideation process fully. |
| Approach: | They propose a visual story ideation task that automates the selection and arrangement of visual assets into coherent sequences that convey expressive storylines. |
| Outcome: | The proposed framework surpasses baseline by 33.5% and 18.5%, respectively, on three metrics. |