Papers by Preethi Seshadri
Crowdsourcing Speech Data for Low-Resource Languages from Low-Income Workers (2020.lrec-1)
Copied to clipboard
Basil Abraham, Danish Goel, Divya Siddarth, Kalika Bali, Manu Chopra, Monojit Choudhury, Pratik Joshi, Preethi Jyoti, Sunayana Sitaram, Vivek Seshadri
| Challenge: | Existing platforms collect labelled speech data from urban speakers whose dialects are often very different from low-income users. |
| Approach: | They propose to collect labelled speech data directly from low-income workers . they collect 109 hours of data from 36 participants in the Marathi language . |
| Outcome: | The proposed approach can provide valuable supplemental earning opportunities to low-income rural and urban workers. |
Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations (2026.acl-long)
Copied to clipboard
| Challenge: | Agentic benchmarks rely on LLM-simulated users to evaluate agent performance . however, the robustness, validity, and fairness of this approach remain unexamined . |
| Approach: | They investigate whether LLM-simulated users are reliable proxies for real human users . they find that agent success rates vary up to 9 percentage points across different LLMs . |
| Outcome: | The results show that simulated users underestimate success on challenging tasks while miscalibrate performance on moderately difficult tasks. |
The Bias Amplification Paradox in Text-to-Image Generation (2024.naacl-long)
Copied to clipboard
| Challenge: | amplification is a phenomenon in which models exacerbate biases or stereotypes in training data. |
| Approach: | They compare gender ratios in training vs. generated images to investigate bias amplification . they find that a model amplifys gender-occupation biases considerably . |
| Outcome: | The proposed model amplifys gender-occupation biases in training data, but it can be attributed to discrepancies between training captions and model prompts. |