Papers by Abhilekh Borah
DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models’ Understanding on Indian Culture (2025.emnlp-main)
Copied to clipboard
Arijit Maji, Raghvendra Kumar, Akash Ghosh, null Anushka, Nemil Shah, Abhilekh Borah, Vanshika Shah, Nishant Mishra, Sriparna Saha
| Challenge: | DRISHTIKON is a first-of-its-kind multimodal and multilingual benchmark centered exclusively on Indian culture. |
| Approach: | They evaluate a wide range of vision-language models across zero-shot and chain-of-thought settings and use them to evaluate cultural understanding of generative AI systems. |
| Outcome: | The DRISHTIKON dataset covers 15 languages, all states and union territories, and incorporating over 64,000 aligned text-image pairs. |
Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations (2025.emnlp-main)
Copied to clipboard
Abhilekh Borah, Chhavi Sharma, Danush Khanna, Utkarsh Bhatt, Gurpreet Singh, Hasnat Md Abdullah, Raghav Kaushik Ravi, Vinija Jain, Jyoti Patel, Shubham Singh, Vasu Sharma, Arpita Vats, Rahul Raja, Aman Chadha, Amitava Das
| Challenge: | a new metric measures the quality of large language models (LLMs) that detects hidden misalignments and jailbreak risks. |
| Approach: | They propose a decoding-invariant metric that measures latent safety failures . they propose 'Alignment Quality Index' to measure latent activations in latent space . |
| Outcome: | The proposed metric detects latent safety failures overlooked by behavioral benchmarks and jailbreaks. |
Don’t Judge a Book by its Cover: Testing LLMs’ Robustness Under Logical Obfuscation (2026.eacl-long)
Copied to clipboard
| Challenge: | obfuscated questions pose significant challenges for large language models . current models parse questions without deep understanding, MIT researchers say . |
| Approach: | They propose a structure-preserving framework for logical obfuscation to test models . they use a logically equivalent framework to obliviate questions to logical equivalents . |
| Outcome: | The proposed framework is a first-of-its-kind diagnostic benchmark with 1,108 questions . obfuscation severely degrades zero-shot performance, the authors show . |