Papers by Deepak Pandita

3 papers
How Many Ratings per Item are Necessary for Reliable Significance Testing? (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods for estimating model reliability are based on a few output responses per item.
Approach: They propose a method to determine whether an existing dataset has enough responses per item to assure reliable null hypothesis statistical testing.
Outcome: The proposed method can help researchers make better decisions about how to collect data for AI evaluation.
Thesis Proposal: Toward a Human-Centered and Perspective-Aware Framework for Reproducible ML Evaluation and AI Alignment (2026.acl-srw)

Copied to clipboard

Challenge: Disagreement arises from subjective human opinion and can vary with one’s identity, beliefs, and social environment.
Approach: They propose a human-centered framework for reproducible ML evaluation and AI alignment that takes disagreement into account when building human-centric AI systems.
Outcome: The proposed framework is based on a human-centered and perspective-aware framework for reproducible ML evaluation and AI alignment.
Rater Cohesion and Quality from a Vicarious Perspective (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent work in reinforcement learning with human feedback (RLHF) highlights the gains in model performance from aligning them to human values.
Approach: They propose to use vicarious annotation to break down disagreement by asking raters how they think others would annotate the data.
Outcome: The proposed method breaks down disagreements by asking raters how they think others would annotate the data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations