Papers by Rijul Magu
Silent Signals, Loud Impact: LLMs for Word-Sense Disambiguation of Coded Dog Whistles (2024.acl-long)
Copied to clipboard
| Challenge: | a dog whistle is a coded communication that carries a secondary meaning to specific audiences and is often weaponized for racial and socioeconomic discrimination. |
| Approach: | They propose an approach for word-sense disambiguation of dog whistles from standard speech using Large Language Models. |
| Outcome: | The proposed method allows disambiguation of dog whistles from standard speech using large language models. |
What About the Scene With the Hitler Reference? HAUNT: A Framework to Probe LLMs’ Self-consistency in Closed Domains Via Adversarial Nudge (2026.acl-long)
Copied to clipboard
| Challenge: | Claude exhibits strong resilience, while GPT and Grok demonstrate moderate resilience . open models fall short significantly, while proprietary models exhibit weak resilience compared to open models . |
| Approach: | They propose a framework for stress testing factual fidelity in large language models in the presence of adversarial nudges. |
| Outcome: | The proposed model is robust to adversarial nudges in two closed domains. |
Auditing LLM Responses to Harmful Stereotypes Targeting Mental Health Groups (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can exhibit imbalanced biases against vulnerable groups, but how they rationalize stereotypes and rights restrictions targeting mental health entities remains underexplored. |
| Approach: | They audit a suite of open-weight LLMs on stereotype-justification prompts tied to mental health identities. |
| Outcome: | The proposed models endorse harmful stereotypes when explicitly asked to justify them, with endorsement varying across model families, versions, and mental health conditions. |