Papers by Chirag Nagpal
Bias in Language Models: Beyond Trick Tests and Towards RUTEd Evaluation (2025.acl-long)
Copied to clipboard
| Challenge: | Standard bias benchmarks are used for large language models to measure the association between social attributes and single-word outputs. |
| Approach: | They adapt three standard bias metrics of next-word prediction to measure gender-occupation bias and develop an analogous RUTEd evaluation in three contexts of real-world LLM use. |
| Outcome: | The proposed benchmarks are robust to lengthening model outputs via a more realistic user prompt in the domain of gender-occupation bias. |