Papers by Deepak Pandita
How Many Ratings per Item are Necessary for Reliable Significance Testing? (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing methods for estimating model reliability are based on a few output responses per item. |
| Approach: | They propose a method to determine whether an existing dataset has enough responses per item to assure reliable null hypothesis statistical testing. |
| Outcome: | The proposed method can help researchers make better decisions about how to collect data for AI evaluation. |
Thesis Proposal: Toward a Human-Centered and Perspective-Aware Framework for Reproducible ML Evaluation and AI Alignment (2026.acl-srw)
Copied to clipboard
| Challenge: | Disagreement arises from subjective human opinion and can vary with one’s identity, beliefs, and social environment. |
| Approach: | They propose a human-centered framework for reproducible ML evaluation and AI alignment that takes disagreement into account when building human-centric AI systems. |
| Outcome: | The proposed framework is based on a human-centered and perspective-aware framework for reproducible ML evaluation and AI alignment. |
Rater Cohesion and Quality from a Vicarious Perspective (2024.findings-emnlp)
Copied to clipboard
Deepak Pandita, Tharindu Cyril Weerasooriya, Sujan Dutta, Sarah Luger, Tharindu Ranasinghe, Ashiqur KhudaBukhsh, Marcos Zampieri, Christopher Homan
| Challenge: | Recent work in reinforcement learning with human feedback (RLHF) highlights the gains in model performance from aligning them to human values. |
| Approach: | They propose to use vicarious annotation to break down disagreement by asking raters how they think others would annotate the data. |
| Outcome: | The proposed method breaks down disagreements by asking raters how they think others would annotate the data. |