Papers by Somdeb Sarkhel

3 papers
Question Modifiers in Visual Question Answering (2022.lrec-1)

Copied to clipboard

Challenge: Visual Question Answering (VQA) is a multi-disciplinary task that requires integration of several key disciplines.
Approach: They develop a model that adds modifiers to questions based on object properties and spatial relationships using Amazon Mechanical Turk data.
Outcome: The proposed model can improve when questions are modified to include more details.
TAME-RD: Text Assisted Replication of Image Multi-Adjustments for Reverse Designing (2024.findings-acl)

Copied to clipboard

Challenge: a new model to reverse design images can be used to replicate image edits on other images based on human instructions in natural language . a study of a dataset of 100K source and edited images shows improvements in accuracy and concordance correlation coefficient .
Approach: They propose a reverse-designing model that automatically learns from image editing operations and natural language instructions to learn fully specified edit operations.
Outcome: The proposed model improves accuracy and concordance correlation scores on multiple datasets.
VISIAR: Empower MLLM for Visual Story Ideation (2025.findings-acl)

Copied to clipboard

Challenge: Existing literature on visual storytelling has not explored the ideation process fully.
Approach: They propose a visual story ideation task that automates the selection and arrangement of visual assets into coherent sequences that convey expressive storylines.
Outcome: The proposed framework surpasses baseline by 33.5% and 18.5%, respectively, on three metrics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations