Papers by Vineet Jain
Comprehensive Multi-Modal Interactions for Referring Image Segmentation (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for RIS compute different forms of interactions sequentially or ignore intra-modal interactions. |
| Approach: | They propose a method which outputs a segmentation map corresponding to the natural language description. |
| Outcome: | The proposed method performs on four benchmark datasets and shows significant performance gains over the existing state-of-the-art methods. |
Scaling Laws and Efficient Inference for Ternary Language Models (2025.acl-long)
Copied to clipboard
Tejas Vaidhya, Ayush Kaushal, Vineet Jain, Francis Couture-Harpin, Prashant Shishodia, Majid Behbahani, Yuriy Nevmyvaka, Irina Rish
| Challenge: | Large language models (LLMs) are increasingly used across research and industry applications, yet their inference efficiency remains a challenge. |
| Approach: | They propose ternary language models that employ quantization-aware training to significantly reduce memory requirements. |
| Outcome: | The proposed ternary language models demonstrate sustained performance gains at scale. |
RiTTA: Modeling Event Relations in Text-to-Audio Generation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing text-to-audio (TTA) generation methods have not explored audio event relation modeling, nor proposed any new framework to enhance this capability. |
| Approach: | They propose a comprehensive relation corpus covering all potential relations in real-world scenarios and a new audio event corpus encompassing commonly heard audios. |
| Outcome: | The proposed framework improves existing models’ relation modeling capability with negligible extra parameters. |