Papers by Ishan Upadhyay

1 papers
GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration (2025.acl-long)

Copied to clipboard

Challenge: Language models are often miscalibrated, leading to confidently incorrect answers.
Approach: They propose a benchmark for language model calibration that incorporates comparison with human calibration.
Outcome: The proposed metric analyzes model calibration errors and identifies types of miscalibration that differ from human behavior.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations