Papers by Kevin Zhu
Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing methods to improve LLM performance have focused on sophisticating the model's step-by-step calculation. |
| Approach: | They propose a question analysis prompting strategy in which the model is prompted to explain the question in 'n' words before solving. |
| Outcome: | The proposed prompt outperforms state-of-the-art prompts on arithmetic and commonsense datasets and consistently ranks among the top-2 prompts. |
Learning Personalized Alignment for Evaluating Open-ended Text Generation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Traditional evaluation metrics rely heavily on lexical similarity with human-written references, showing poor correlation with human judgments and failing to account for alignment with the diversity of human preferences. |
| Approach: | They propose an interpretable evaluation framework that evaluates alignment with specific human preferences by providing detailed comments and fine-grained scoring. |
| Outcome: | The proposed framework outperforms GPT-4 in Kendall correlation and accuracy with zero-shot reviewers. |
MiSCHiEF: A Benchmark in Minimal-Pairs of Safety and Culture for Holistic Evaluation of Fine-Grained Image-Caption Alignment (2026.eacl-short)
Copied to clipboard
Sagarika Banerjee, Tangatar Madi, Advait Swaminathan, Jolie Nguyen, Shivank Garg, Kevin Zhu, Vasu Sharma
| Challenge: | Fine-grained image-caption alignment is crucial for vision-language models in socially critical contexts. |
| Approach: | They present a benchmarking dataset for fine-grained image-caption alignment in safety and culture contexts. |
| Outcome: | The proposed benchmarks show that models perform better at confirming correct pairs than rejecting incorrect ones on dual alignment tasks. |
A Dashboard for Mitigating the COVID-19 Misinfodemic (2021.eacl-demos)
Copied to clipboard
Zhengyuan Zhu, Kevin Meng, Josue Caraballo, Israa Jaradat, Xiao Shi, Zeyu Zhang, Farahnaz Akrami, Haojin Liao, Fatma Arslan, Damian Jimenez, Mohanmmed Samiul Saeef, Paras Pathak, Chengkai Li
| Challenge: | a new public dashboard aims to understand the impact of the COVID-19 misinfodemic on Twitter . the dashboard uses a curated catalog of COVId-19 related facts and debunks of misinformation . |
| Approach: | They propose a public dashboard that matches tweets with COVID-19 misinformation . they also propose experiments to analyze the spread of misinformation on twitter . |
| Outcome: | The proposed dashboard uses a curated catalog of COVID-19 related facts and debunks misinformation . it shows the most prevalent information from the catalog among Twitter users in user-selected geographic regions . |
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing LLM leaderboards often reference scores reported in other papers without consistent settings and prompts, which may encourage cherry-picking favored settings and for better results. |
| Approach: | They propose an open-source and reproducible LLM evaluation suite built on top of OpenAI Evals that systematically evaluates 10+ leading LLMs and OpenAI’s legacy models on 20+ curated benchmarks across 7 capability categories. |
| Outcome: | The evaluation suite is built on top of OpenAI Evals and evaluates 10+ leading LLMs and OpenAI’s legacy models on 20+ curated benchmarks across 7 capability categories. |
VIEWS: Entity-Aware News Video Captioning (2024.emnlp-main)
Copied to clipboard
Hammad Ayyubi, Tianqi Liu, Arsha Nagrani, Xudong Lin, Mingda Zhang, Anurag Arnab, Feng Han, Yukun Zhu, Xuande Feng, Kevin Zhang, Jialu Liu, Shih-Fu Chang
| Challenge: | Existing video captioning benchmarks and models produce generic captions for videos that lack specific identification of individuals, locations, or organizations. |
| Approach: | They propose a task of directly summarizing news videos into captions that are entity-aware . they validate the effectiveness of their approach across three video captioning models . |
| Outcome: | The proposed approach is effective across three video captioning models. |
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts (2025.emnlp-main)
Copied to clipboard
| Challenge: | Current large language models struggle to answer questions that span tens of thousands of tokens. |
| Approach: | They evaluate 1–4 hop QA over 64k–128k-token excerpts from 83 novels . they find consistent accuracy drops with increased hops and context length . |
| Outcome: | The novelhopqa benchmark evaluates 1–4 hop QA over 64k–128k-token excerpts from 83 public-domain novels. |
EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing benchmarks often overlook intra-language variations, leaving speakers of non-standard dialects underserved. |
| Approach: | EnDive evaluates seven state-of-the-art large language models across tasks . human evaluations confirm high translation quality, with average scores of at least 6.02/7 . |
| Outcome: | EnDive evaluates state-of-the-art large language models across language understanding, reasoning, mathematics, logic tasks. |
An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language Models (2024.findings-acl)
Copied to clipboard
Gantavya Bhatt, Yifang Chen, Arnav Das, Jifan Zhang, Sang Truong, Stephen Mussmann, Yinglun Zhu, Jeff Bilmes, Simon Du, Kevin Jamieson, Jordan Ash, Robert Nowak
| Challenge: | Supervised finetuning (SFT) on instruction datasets has shown immense potential in improving the zero-shot generalization capabilities observed in large language models (LLMs). |
| Approach: | They propose to use experimental design to minimize the computational cost of active learning by identifying useful subsets of samples to annotate from an unlabeled pool. |
| Outcome: | The proposed methods save 50% of the annotation cost compared to random sampling on generative tasks. |
Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs (2026.acl-long)
Copied to clipboard
Nikita Afonin, Nikita Andriianov, Vahagn Hovhannisyan, Nikhil Bageshpura, Kyle Liu, Kevin Zhu, Sunishchal Dev, Ashwinee Panda, Oleg Rogov, Elena Tutubalina, Alexander Panchenko, Mikhail Seleznyov
| Challenge: | Recent studies have documented emergent misalignment in language models adapted on narrow examples . Emergent misalignement occurs when models are trained on narrow set of misallocated examples resulting in harmful or misleading responses . |
| Approach: | They propose to explain in-context EM as conflict between safety objectives and context-following behavior. |
| Outcome: | The proposed model is adapted on 16 in-context examples and produces misaligned responses to benign queries. |
Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning (2025.naacl-srw)
Copied to clipboard
Shaun Lee Baek, Shaun Esua-Mensah, Cyrus Tsui, Sejan Vigneswaralingam, Abdullah Alali, Michael Lu, Vasu Sharma, Kevin Zhu
| Challenge: | Large Language Models (LLMs) are primarily trained on high-resource natural languages, limiting their effectiveness in low-resourced settings and in tasks requiring deep logical reasoning. |
| Approach: | They propose to use a dataset of logical propositions from Lean into a custom logical language to evaluate LLMs' logical reasoning and generalization capabilities in a controlled environment. |
| Outcome: | The proposed model improves accuracy and accuracy beyond 20,000 training samples. |