Papers by Vipul Gupta

7 papers
LLMs Assist NLP Researchers: Critique Paper (Meta-)Reviewing (2024.emnlp-main)

Copied to clipboard

Challenge: a comparative analysis of paper (meta-)reviews by large language models (LLMs) aims to identify and distinguish LLMs from human activities .
Approach: They present a comparative analysis to identify and distinguish LLM activities from human activities.
Outcome: The proposed analysis aims to improve recognition of instances when someone implicitly uses LLMs for reviewing activities.
Improving Model Evaluation using SMART Filtering of Benchmark Datasets (2025.naacl-long)

Copied to clipboard

Challenge: Creating high quality human-annotated datasets is difficult due to dataset saturation.
Approach: They propose a method to filter a subset of test examples from existing benchmarks by removing less informative and lower quality examples.
Outcome: The proposed method reduces dataset size by 48% while increasing Pearson correlation with rankings from ChatBot Arena.
An Audit on the Perspectives and Challenges of Hallucinations in NLP (2024.emnlp-main)

Copied to clipboard

Challenge: 103 peer-reviewed publications on hallucination in large language models (LLMs) are characterized by a lack of agreement with the term ‘hallucination’ in the field of NLP.
Approach: They examine 103 peer-reviewed publications on hallucination in large language models (LLMs) and conduct a survey with 171 practitioners from the field of NLP and AI to capture varying perspectives on halllucination.
Outcome: The findings highlight the need for explicit definitions and frameworks outlining hallucination within NLP and highlight potential challenges.
PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Frontier models often lack a view of performance on open-ended, economically consequential tasks in high-stakes professional domains where practical returns matter most.
Approach: They introduce a professional reasoning benchmark that recruits 182 qualified professionals to contribute questions inspired by their workflows.
Outcome: The proposed model outperforms other models in 114 countries and 47 US jurisdictions on hard subsets.
Development and Benchmarking of a Blended Human-AI Qualitative Research Assistant (2026.acl-industry)

Copied to clipboard

Challenge: Qualitative research emphasizes constructing meaning through iterative engagement with textual data.
Approach: They present and benchmark a qualitative research assistant system that allows researchers to identify themes and annotate datasets.
Outcome: The proposed system achieves an inter-rater reliability between Muse and humans of Cohen’s = 0.7 for well-specified codes.
The Sentiment Problem: A Critical Survey towards Deconstructing Sentiment Analysis (2023.emnlp-main)

Copied to clipboard

Challenge: Existing research reveals a notable absence of interdisciplinary endeavors to comprehend the social dimensions of sentiment analysis, encompassing aspects like emotion and fairness.
Approach: They propose an ethics sheet encompassing critical inquiries to guide practitioners in ensuring equitable utilization of SA.
Outcome: The proposed ethics sheet outlines the importance of adopting an interdisciplinary approach to defining sentiment in SA and offers a pragmatic solution for its implementation.
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks (2026.acl-long)

Copied to clipboard

Challenge: Multiple-choice question answering (MCQ) is standard in NLP, but benchmarks lack rigorous quality control.
Approach: They propose an education-inspired toolkit that uses LLM judges to flag flaws in MCQs . they validate the tool with annotations and run it to audit 12 benchmarks based on 19-rule education rubric .
Outcome: The proposed toolkit flags three common MCQ flaws based on a 19-rule education rubric . contaminated MCqs tend to inflate accuracy, while writing errors lower it and change rankings .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations