Papers by Vishwa Vinay
Robustness of Fusion-based Multimodal Classifiers to Cross-Modal Content Dilutions (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing work has focused on understanding the robustness of vision-and-language models to imperceptible variations on benchmark tasks. |
| Approach: | They develop a model that generates additional dilution text that maintains relevance and topical coherence with the image and existing text, and when added to the original text, leads to misclassification of the multimodal input. |
| Outcome: | The proposed model outperforms fusion-based classifiers on Crisis Humanitarianism and Sentiment Detection tasks by 23.3% and 22.5% in presence of dilutions generated by the model. |