Papers by Ishan Upadhyay
GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration (2025.acl-long)
Copied to clipboard
| Challenge: | Language models are often miscalibrated, leading to confidently incorrect answers. |
| Approach: | They propose a benchmark for language model calibration that incorporates comparison with human calibration. |
| Outcome: | The proposed metric analyzes model calibration errors and identifies types of miscalibration that differ from human behavior. |