Papers by Tushar Vatsal
Evaluating Concurrent Robustness of Language Models Across Diverse Challenge Sets (2024.emnlp-main)
Copied to clipboard
| Challenge: | Language models display sensitivity to input perturbations, causing concerns about trust among users. |
| Approach: | They propose a methodology to examine how input perturbations affect language models across various scales, including pre-trained models and large language models. |
| Outcome: | The proposed methods enhance the model’s robustness to input perturbations and if exposure to one perturbation enhances or diminishes its performance with respect to other perturbations. |
Automated Digitization of Unstructured Medical Prescriptions (2023.acl-industry)
Copied to clipboard
| Challenge: | e-commerce prescription ordering is challenging in emerging markets since prescriptions are paper-based, unstructured and often, handwritten. |
| Approach: | They propose a prescription digitization system for online medicine ordering built with minimal supervision. |
| Outcome: | The proposed system achieves +5.9% gain in precision@3 and +5.6% in recall@3 over baselines on medication attribute extraction. |
NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent advances in large language models have demonstrated their strong performance on IQ test questions, achieving high scores across many languages. |
| Approach: | They propose a dataset to evaluate cognitive multimodal reasoning and problem-solving skills of large models. |
| Outcome: | The proposed dataset contains 2,728 multiple-choice questions and 4,642 images spanning 26 categories. |