Papers by Bhavik Chandna
A Counterfactual Explanation Framework for Retrieval Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing literature on explainability of information retrieval has focused on illustrating the concept of relevance concerning a retrieval model. |
| Approach: | They propose to add terms to a document to improve its ranking to answer the question of which words played a role in not being favored by a retrieval model. |
| Outcome: | The proposed framework predicts counterfactuals for statistical and deep-learning models. |
ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing datasets for evaluating LMM robustness lack exploration of extremist content . existing models lack diverse image generation models and comprehensive coverage of historical events . |
| Approach: | They propose a benchmark dataset to assess LMM models against extremist content . ExtremeAIGC simulates real-world events and malicious use cases . |
| Outcome: | a new benchmark dataset and evaluation framework assesses LMM models against extremist content. |
XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing safety evaluations rely on binary labels, overlooking the nuanced risk these outputs pose. |
| Approach: | They propose a framework to assess the severity of extremist content generated by Large Language Models (LLMs) it categorizes model responses into five danger levels (0–4) defined by degree of extremism endorsement . |
| Outcome: | The proposed framework categorizes model responses into five danger levels (0–4) defined by degree of extremist endorsement, enabling nuanced analysis of failure frequency and severity. |