Papers by Raghuveer Rao
X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing text-to-video retrieval systems use embedding models for feature extraction and compute cosine similarities for ranking. |
| Approach: | They propose an explainable retrieval framework upon LLM CoT reasoning to replace embedding models for feature extraction and ranking. |
| Outcome: | The proposed retrieval framework improves retrieval performance and produces detailed rationales. |
MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper (2025.emnlp-main)
Copied to clipboard
Runjia Zeng, Guangyan Sun, Qifan Wang, Tong Geng, Sohail Dianat, Xiaotian Han, Raghuveer Rao, Xueling Zhang, Cheng Han, Lifu Huang, Dongfang Liu
| Challenge: | Empirical evaluations show that Mixture of Expert Prompt Tuning outperforms state-of-the-art parameter efficient baselines on SuperGLUE. |
| Approach: | They propose a pretrain-then-fine-tune paradigm for manifold mapping using multiple prompt experts. |
| Outcome: | Empirical results show that the proposed approach outperforms state-of-the-art methods on SuperGLUE while reducing activated prompts by 79.25%. |
Visual Self-Refinement for Autoregressive Models (2025.findings-emnlp)
Copied to clipboard
Jiamian Wang, Ziqi Zhou, Chaithanya Kumar Mummadi, Sohail Dianat, Majid Rabbani, Raghuveer Rao, Chen Qiu, Zhiqiang Tao
| Challenge: | Autoregressive models excel in sequential modeling but the spatial nature of visual signals conflicts with the sequential dependencies of next-token prediction, leading to suboptimal results. |
| Approach: | They propose a plug-and-play refinement module to enhance the spatial correspondence modeling within the generated visual sequence. |
| Outcome: | The proposed module enhances vision-language modeling under a shared sequential prediction framework. |