Papers by Sangmitra Madhusudan

4 papers
Fine-Tuned LLMs are “Time Capsules” for Tracking Societal Bias Through Books (2025.naacl-long)

Copied to clipboard

Challenge: We develop a corpus comprising 593 fictional books across seven decades (1950-2019) to track bias evolution.
Approach: They develop a method to trace and quantify bias evolution using fine-tuned LLMs on fictional books across seven decades to track bias evolution.
Outcome: The proposed method traces and quantifies bias evolution in a corpus of 593 fictional books across seven decades.
STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models that assess explicit and implicit biases are based on a single scenario . a dataset of 450 offensive progressions contains 2,700 sentences of varying severity .
Approach: They evaluate a dataset of offensive progressions that contain 2,700 sentences . they find that even the best-performing models detect bias inconsistently .
Outcome: The proposed dataset shows that even the best-performing models detect bias inconsistently . aligning models with human judgments on STOP can improve answer rates on sensitive tasks by 191% .
Common to Whom? Regional Cultural Commonsense and LLM Bias in India (2026.acl-long)

Copied to clipboard

Challenge: Existing cultural commonsense benchmarks treat nations as monolithic, assuming uniform practices within national boundaries.
Approach: They evaluate eight state-of-the-art LLMs and find two critical gaps . commonsense knowledge is fundamentally long-tailed, with most facts rare in training data .
Outcome: The proposed model achieves only 13.4%–20.9% accuracy on region-specific questions and exhibits geographic bias over-selecting Central and North India as the "default" while under-representing East and West.
The Dog the Cat Chased Stumped the Model: Measuring When Language Models Abandon Structure for Shortcuts (2026.eacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) can explain quantum mechanics and write sophisticated code, yet fail to parse sentences like "The cat that the mouse feared chased meowed"
Approach: They propose a framework to distinguish structural understanding from semantic pattern matching . they use a set of 9,720 comprehension questions on center-embedded sentences .
Outcome: a new framework shows that models lose performance when they abandon structural analysis for semantic associations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations