Papers by Vipul Gupta
LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing (2024.emnlp-main)
Copied to clipboard
Jiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng, Shuaiqi Liu, Renze Lou, Henry Zou, Pranav Narayanan Venkit, Nan Zhang, Mukund Srinath, Haoran Zhang, Vipul Gupta, Yinghui Li, Tao Li, Fei Wang, Qin Liu, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang, Ying Su, Raj Shah, Ruohao Guo, Jing Gu, Haoran Li, Kangda Wei, Zihao Wang, Lu Cheng, Surangika Ranathunga, Meng Fang, Jie Fu, Fei Liu, Ruihong Huang, Eduardo Blanco, Yixin Cao, Rui Zhang, Philip Yu, Wenpeng Yin
| Challenge: | a comparative analysis of paper (meta-)reviews by large language models (LLMs) aims to identify and distinguish LLMs from human activities . |
| Approach: | They present a comparative analysis to identify and distinguish LLM activities from human activities. |
| Outcome: | The proposed analysis aims to improve recognition of instances when someone implicitly uses LLMs for reviewing activities. |
Improving Model Evaluation using SMART Filtering of Benchmark Datasets (2025.naacl-long)
Copied to clipboard
| Challenge: | Creating high quality human-annotated datasets is difficult due to dataset saturation. |
| Approach: | They propose a method to filter a subset of test examples from existing benchmarks by removing less informative and lower quality examples. |
| Outcome: | The proposed method reduces dataset size by 48% while increasing Pearson correlation with rankings from ChatBot Arena. |
An Audit on the Perspectives and Challenges of Hallucinations in NLP (2024.emnlp-main)
Copied to clipboard
Pranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs, Mukund Srinath, Koustava Goswami, Sarah Rajtmajer, Shomir Wilson
| Challenge: | 103 peer-reviewed publications on hallucination in large language models (LLMs) are characterized by a lack of agreement with the term ‘hallucination’ in the field of NLP. |
| Approach: | They examine 103 peer-reviewed publications on hallucination in large language models (LLMs) and conduct a survey with 171 practitioners from the field of NLP and AI to capture varying perspectives on halllucination. |
| Outcome: | The findings highlight the need for explicit definitions and frameworks outlining hallucination within NLP and highlight potential challenges. |
PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning (2026.acl-long)
Copied to clipboard
Afra Feyza Akyürek, Advait Gosai, Chen Bo Calvin Zhang, Vipul Gupta, Jaehwan Jeong, Anisha Gunjal, Tahseen Rabbani, Maria Mazzone, David Randolph IV, Mohammad Mahmoudi Meymand, Gurshaan Chattha, Paula Rodriguez, Diego A. Mares Buendia, Pavit Singh, Michael Liu, Subodh Chawla, Peter Cline, Lucy Ogaz, Ernesto Gabriel Hernández Montoya, Zihao Wang, Pavi Bhatter, Marcos Ayestaran, Bing Liu, Yunzhong He
| Challenge: | Frontier models often lack a view of performance on open-ended, economically consequential tasks in high-stakes professional domains where practical returns matter most. |
| Approach: | They introduce a professional reasoning benchmark that recruits 182 qualified professionals to contribute questions inspired by their workflows. |
| Outcome: | The proposed model outperforms other models in 114 countries and 47 US jurisdictions on hard subsets. |
Development and Benchmarking of a Blended Human-AI Qualitative Research Assistant (2026.acl-industry)
Copied to clipboard
Joseph Matveyenko, James Liu, John David Parsons, Ryan Brown, Alina I. Palimaru, Vipul Gupta, Prateek Puri
| Challenge: | Qualitative research emphasizes constructing meaning through iterative engagement with textual data. |
| Approach: | They present and benchmark a qualitative research assistant system that allows researchers to identify themes and annotate datasets. |
| Outcome: | The proposed system achieves an inter-rater reliability between Muse and humans of Cohen’s = 0.7 for well-specified codes. |
The Sentiment Problem: A Critical Survey towards Deconstructing Sentiment Analysis (2023.emnlp-main)
Copied to clipboard
Pranav Venkit, Mukund Srinath, Sanjana Gautam, Saranya Venkatraman, Vipul Gupta, Rebecca Passonneau, Shomir Wilson
| Challenge: | Existing research reveals a notable absence of interdisciplinary endeavors to comprehend the social dimensions of sentiment analysis, encompassing aspects like emotion and fairness. |
| Approach: | They propose an ethics sheet encompassing critical inquiries to guide practitioners in ensuring equitable utilization of SA. |
| Outcome: | The proposed ethics sheet outlines the importance of adopting an interdisciplinary approach to defining sentiment in SA and offers a pragmatic solution for its implementation. |
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks (2026.acl-long)
Copied to clipboard
Nishant Balepur, Bhavya Rajasekaran, Hyunjin Jane Oh, Michael Xie, Atrey Desai, Vipul Gupta, Steven James Moore, Eunsol Choi, Rachel Rudinger, Jordan Lee Boyd-Graber
| Challenge: | Multiple-choice question answering (MCQ) is standard in NLP, but benchmarks lack rigorous quality control. |
| Approach: | They propose an education-inspired toolkit that uses LLM judges to flag flaws in MCQs . they validate the tool with annotations and run it to audit 12 benchmarks based on 19-rule education rubric . |
| Outcome: | The proposed toolkit flags three common MCQ flaws based on a 19-rule education rubric . contaminated MCqs tend to inflate accuracy, while writing errors lower it and change rankings . |