Papers by Kanishk Jain

2 papers
Comprehensive Multi-Modal Interactions for Referring Image Segmentation (2022.findings-acl)

Copied to clipboard

Challenge: Existing methods for RIS compute different forms of interactions sequentially or ignore intra-modal interactions.
Approach: They propose a method which outputs a segmentation map corresponding to the natural language description.
Outcome: The proposed method performs on four benchmark datasets and shows significant performance gains over the existing state-of-the-art methods.
Benchmarking Vision Language Models for Cultural Understanding (2024.emnlp-main)

Copied to clipboard

Challenge: Recent multimodal vision-language models have shown impressive performance in tasks such as image-to-text generation, visual question answering, and image captioning.
Approach: They propose a visual question-answering benchmark to assess VLMs' cultural understanding of various facets of culture from 11 countries across 5 continents.
Outcome: The visual question-answering benchmark aims to assess VLMs' cultural understanding across regions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations