Papers by Yongmin Kim
Flexible Visual Grounding (2022.acl-srw)
Copied to clipboard
| Challenge: | Existing visual grounding datasets require queries to be answerable, but in multimedia data, many entities cannot be grounded to the image, resulting in unanswerable visual ground. |
| Approach: | They propose a method to ground to a pseudo image region for unanswerable queries . they add a query that cannot be grounded to the image and train it to ground . |
| Outcome: | The proposed model can handle answerable and unanswerable visual grounding with high accuracy on the proposed datasets. |