When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing static benchmarks that measure task performance often rely on a simple input-output configuration. |
| Approach: | They propose an evaluation pipeline that evaluates code models with different feedback types in an interactive setting. |
| Outcome: | The proposed evaluation pipeline compares model-user collaboration with static benchmarks by obfuscating inputs to a simulated user. |
Similar Papers
CodeArena: A Collective Evaluation Platform for LLM Code Generation (2025.acl-demo)
Copied to clipboard
Mingzhe Du, Anh Tuan Luu, Bin Ji, Xiaobao Wu, Yuhao Qing, Dong Huang, Terry Yue Zhuo, Qian Liu, See-Kiong Ng
| Challenge: | Large Language Models (LLMs) have reshaped code generation, but persistent challenges impede accurate assessment. |
| Approach: | They propose an online evaluation framework tailored for large language models to assess their coding capabilities. |
| Outcome: | a new evaluation framework for large language models (LLMs) provides unbiased, unbiased evaluations and open access to solutions and test cases. |
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used to judge code, but their reliability remains poorly understood. |
| Approach: | They propose a benchmark to evaluate Large Language Models as code judges . they find that small reasoning models outperform larger non-reasoning models . |
| Outcome: | The proposed benchmark evaluates LLM-as-a-Judge models across three coding tasks. |
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent work on LLM-as-a-Judge has reported higher correlations with human judgments due to its static nature. |
| Approach: | They propose a framework that leverages multi-turn interactions where the LLM interviewer actively provides feedback on responses and poses follow-up questions to the evaluated LLM. |
| Outcome: | The proposed framework evaluates six models on reasoning, factuality and instruction-following tasks. |
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation (2026.findings-acl)
Copied to clipboard
Zhichao Shi, Xuhui Jiang, Chengjin Xu, Cangli Yao, Shengjie Ma, Yinghan Shen, Zixuan Li, Jian Guo, Yuanzhuo Wang
| Challenge: | Current evaluation methods for large language models rely on static benchmarks . limited knowledge coverage and fixed difficulties hinder the targeted optimizations resulting in superficial evaluations of LLMs - a problem that has been addressed by JudgeAgent . |
| Approach: | They propose a knowledge-driven and dynamic evaluation framework for large language models . judgeAgent leverages LLM agents equipped with context graphs to traverse knowledge structures . |
| Outcome: | The proposed framework can achieve comprehensive evaluations and facilitate effective model iterations. |
KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models (2024.acl-long)
Copied to clipboard
Zhuohao Yu, Chang Gao, Wenjin Yao, Yidong Wang, Wei Ye, Jindong Wang, Xing Xie, Yue Zhang, Shikun Zhang
| Challenge: | Existing methods to detect contaminated texts focus on quantifying contamination status instead of accurately gauging model performance. |
| Approach: | They propose a Knowledge-grounded Interactive Evaluation framework which incorporates an LLM-powered “interactor” role for the first time to accomplish a dynamic contamination-resilient evaluation. |
| Outcome: | The proposed framework is based on a question in a standard LLM benchmark and can be used to evaluate models in real-world conversations. |
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks lack the ability to automatically evaluate from users’ perspective and lack the explainability of the results of LLM agents’ code generation capabilities. |
| Approach: | They propose a new benchmark for LLM agents' automated evaluation by simulating user interaction. |
| Outcome: | The proposed benchmark can evaluate the generated projects by user interaction simulation and by code similarity through existing objective indicators. |
Beyond Static Evaluation: A Dynamic Approach to Assessing AI Assistants’ API Invocation Capabilities (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing evaluation methods for human-machine interactions are static and can be misleading. |
| Approach: | They propose to use a LLM-based user agent to assess an assistant's API call capability without human involvement. |
| Outcome: | The proposed method mirrors real human conversation patterns in human-machine interactions, and shows that it aligns more closely with human assessment. |
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development (2026.findings-acl)
Copied to clipboard
Jie Yang, Honglin Guo, Li Ji, Jiazheng Zhou, Rui Zheng, Zhikai Lei, Shuo Zhang, Zhiheng Xi, Shichun Liu, Yuxin Wang, Bo Wang, Yining Zheng, Tao Gui, Xipeng Qiu
| Challenge: | Large Language Models (LLMs) have redefined the role of AI in software engineering . current benchmarks focus on localized code generation, but neglect dynamic, full-process requirements of real-world engineering. |
| Approach: | They propose a benchmark to evaluate agentic backend coding within a realistic, executable workflow. |
| Outcome: | The ABC-Bench benchmark evaluates agentic backend coding within a realistic, executable workflow. |
LLMeBench: A Flexible Framework for Accelerating LLMs Benchmarking (2024.eacl-demo)
Copied to clipboard
Fahim Dalvi, Maram Hasanain, Sabri Boughorbel, Basel Mousi, Samir Abdaljalil, Nizi Nazar, Ahmed Abdelali, Shammur Absar Chowdhury, Hamdy Mubarak, Ahmed Ali
| Challenge: | Recent development and success of Large Language Models necessitate evaluation of their performance across diverse NLP tasks in different languages. |
| Approach: | They propose a framework that can be customized to evaluate LLMs for any NLP task, regardless of language. |
| Outcome: | The LLMeBench framework can be customized to evaluate LLMs for any NLP task, regardless of language. |
FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs’ Responsiveness to Human Feedback (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing research focuses on benchmarking LLMs in single-turn dialogues, neglecting the nuanced nature of human feedback within real-world usage scenarios. |
| Approach: | They propose a fine-grained, multi-task benchmark designed to evaluate LLMs’ responsiveness to human feedback under real-world usage scenarios in Chinese. |
| Outcome: | The proposed benchmarks show that human feedback can significantly impact LLMs’ responsiveness in real-world usage scenarios. |