Papers by Sai Rajeswar
StarFlow: Generating Structured Workflow Outputs From Sketch Images (2026.eacl-long)
Copied to clipboard
Patrice Bechard, Chao Wang, Amirhossein Abaskohi, Juan A. Rodriguez, Christopher Pal, David Vazquez, Spandana Gella, Sai Rajeswar, Perouz Taslakian
| Challenge: | Despite being widely used, building workflows can be complex, often requiring manual configuration through low-code platforms or visual programming tools. |
| Approach: | They propose a framework for generating structured workflow outputs from sketches using vision-language models to automate the process. |
| Outcome: | The proposed framework outperforms large vision-language models in the task of generating structured workflow outputs from sketches and diagrams. |
Grammar Search for Multi-Agent Systems (2026.acl-long)
Copied to clipboard
Mayank Singh, Vikas Yadav, Shiva Krishna Reddy Malay, Shravan Nayak, Sai Rajeswar, Sathwik Tejaswi Madhusudhan, Eduardo Blanco
| Challenge: | Several prior approaches have relied on LLM-based free-form search over the code space. |
| Approach: | They propose a more structured framework that explores the same space through a fixed set of composable components. |
| Outcome: | The proposed framework outperforms existing approaches on most benchmarks across two backbone LLMs and two domains: mathematics and question answering. |
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval (2025.emnlp-industry)
Copied to clipboard
Ahmed Masry, Megh Thakkar, Patrice Bechard, Sathwik Tejaswi Madhusudhan, Rabiul Awal, Shambhavi Mishra, Akshay Kalkunte Suresh, Srivatsava Daruru, Enamul Hoque, Spandana Gella, Torsten Scholak, Sai Rajeswar
| Challenge: | Existing methods for multimodal document retrieval often replicate techniques developed for text-only retrieval. |
| Approach: | They propose a document retrieval model that bridges the gap between multimodal representation learning and document retrievals by providing external knowledge as context. |
| Outcome: | The proposed model achieves 3.61% improvement over existing retrieval models on the ViDoRe V2 benchmark, showing stronger generalization to out-of-domain benchmarks. |
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation (2025.emnlp-main)
Copied to clipboard
Rabiul Awal, Mahsa Massoud, Aarash Feizi, Zichao Li, Suyuchen Wang, Christopher Pal, Aishwarya Agrawal, David Vazquez, Siva Reddy, Juan A. Rodriguez, Perouz Taslakian, Spandana Gella, Sai Rajeswar
| Challenge: | Existing benchmarks focus on specific aspects of web tasks but lack comprehensive coverage. |
| Approach: | They propose a multilingual benchmark that evaluates three core web tasks: (1) website visual question answering, (2) code editing involving HTML/CSS/JavaScript, and (3) mockup-to-code generation. |
| Outcome: | The proposed model performs well on basic information extraction, but struggles with reasoning and grounding, editing code to preserve functionality, and generating design-to-code that maintains hierarchy and supports multilingual content. |