UniDebugger: Hierarchical Multi-Agent Framework for Unified Software Debugging (2025.emnlp-main)
Copied to clipboard
Cheryl Lee, Chunqiu Steven Xia, Longji Yang, Jen-tse Huang, Zhouruixing Zhu, Lingming Zhang, Michael R. Lyu
| Challenge: | Existing LLMs focus on isolated steps and struggle with complex bugs. |
| Approach: | They propose a framework for unified debugging through multi-agent synergy . it mimics the entire cognitive processes of developers with each agent specialized as a particular component of this process . |
| Outcome: | The proposed framework outperforms state-of-the-art methods on repo-level benchmarks. |
Similar Papers
COAST: Enhancing the Code Debugging Ability of LLMs through Communicative Agent Based Data Synthesis (2025.findings-naacl)
Copied to clipboard
Weiqing Yang, Hanbin Wang, Zhenghao Liu, Xinze Li, Yukun Yan, Shuo Wang, Yu Gu, Minghe Yu, Zhiyuan Liu, Ge Yu
| Challenge: | Existing code debugging benchmarks focus on the Code Repair stage of the code generation process. |
| Approach: | They propose a framework to evaluate the debugging abilities of large language models by emulating the human debug process. |
| Outcome: | The proposed framework outperforms human-curated and GPT-4-generated training data, enabling 7B-scale LLMs to achieve comparable debugging performance to GPT-3.5. |
RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models (2025.findings-emnlp)
Copied to clipboard
Jingjing Liu, Zeming Liu, Zihao Cheng, Mengliang He, Xiaoming Shi, Yuhang Guo, Xiangrong Zhu, Yuanfang Guo, Yunhong Wang, Haifeng Wang
| Challenge: | Large Language Models (LLMs) have exhibited significant proficiency in code debugging, especially in automatic program repair. |
| Approach: | They propose a repository-level code debugging dataset with 22 subtypes of errors that supports 8 commonly used programming languages and 3 debug tasks. |
| Outcome: | The proposed dataset supports 8 commonly used programming languages and 3 debugging tasks. |
AUTOGEN STUDIO: A No-Code Developer Tool for Building and Debugging Multi-Agent Systems (2024.emnlp-demo)
Copied to clipboard
Victor Dibia, Jingya Chen, Gagan Bansal, Suff Syed, Adam Fourney, Erkang Zhu, Chi Wang, Saleema Amershi
| Challenge: | Multi-agent systems are emerging as effective pattern for solving long-running, complex tasks in numerous do- mains. |
| Approach: | They propose a no-code developer tool for rapidly prototyping, debugging, and evaluating multi-agent work flows built upon the AUTOGEN framework. |
| Outcome: | The proposed tool provides an intuitive drag-and-drop UI for agent workflow specification, interactive evaluation and debugging of workflows, and a gallery of reusable agent components. |
CodeSim: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debugging (2025.findings-naacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have made significant strides in code generation and problem solving. |
| Approach: | They propose a multi-agent code generation framework that integrates human-like perception to address the stages of program synthesis. |
| Outcome: | The proposed framework achieves state-of-the-art (pass@1) results and shows potential for even greater enhancement when cascaded with external debuggers. |
Towards Self-Improving Error Diagnosis in Multi-Agent Systems (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing diagnostic approaches rely on expensive expert annotations and ”LLM-as-a-judge” paradigms. |
| Approach: | They propose a framework for semantic failure attribution that identifies responsible agents and the originating error step. |
| Outcome: | The proposed framework outperforms baselines in step-level localization and validation. |
DebugBench: Evaluating Debugging Capability of Large Language Models (2024.findings-acl)
Copied to clipboard
Runchu Tian, Yining Ye, Yujia Qin, Xin Cong, Yankai Lin, Yinxu Pan, Yesai Wu, Hui Haotian, Liu Weichuan, Zhiyuan Liu, Maosong Sun
| Challenge: | Large language models (LLMs) have demonstrated exceptional coding capabilities, but their debugging capabilities remain relatively unexplored. |
| Approach: | They propose a debugging benchmark consisting of 4,253 LLMs with four major bug categories and 18 minor types in C++, Java, and Python. |
| Outcome: | The proposed benchmark covers four major bug categories and 18 minor types in C++, Java, and Python. |
Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step by Step (2024.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are leading progress in code generation, but they are underutilized in the literature. |
| Approach: | They propose a debugging framework that allows LLMs to refine their generated programs with the runtime execution information. |
| Outcome: | The proposed framework improves the baseline performance by 9.8% across the HumanEval, MBPP, and TransCoder benchmarks. |
Instruct, Not Assist: LLM-based Multi-Turn Planning and Hierarchical Questioning for Socratic Code Debugging (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Current large language models often give away solutions directly, making them ineffective instructors. |
| Approach: | They propose to use a state space-based planning algorithm to build a question tree based on a student's knowledge state to help students independently identify and resolve errors. |
| Outcome: | The proposed model is able to debug code efficiently with minimal turns and highly Socratic questioning. |
AutoDetect: Towards a Unified Framework for Automated Weakness Detection in Large Language Models (2024.findings-emnlp)
Copied to clipboard
Jiale Cheng, Yida Lu, Xiaotao Gu, Pei Ke, Xiao Liu, Yuxiao Dong, Hongning Wang, Jie Tang, Minlie Huang
| Challenge: | Large Language Models (LLMs) exhibit significant but subtle weaknesses, such as mistakes in instruction-following or coding tasks. |
| Approach: | They propose a framework to automatically expose weaknesses in Large Language Models (LLMs) they use three LLM-powered agents to perform comprehensive weakness identification . |
| Outcome: | The proposed framework shows that it is more effective than untargeted data augmentation methods like Self-Instruct to identify weaknesses in LLMs. |
PROTEA: Offline Evaluation and Iterative Refinement for Multi-Agent LLM Workflows (2026.acl-demo)
Copied to clipboard
| Challenge: | Multi-agent LLM workflows are notoriously difficult to debug and refine. |
| Approach: | They propose a unified UI that closes the loop for offline, test-case–driven improvement of multi-agent LLM workflows. |
| Outcome: | PROTEA performs backward node evaluation and proposes a targeted prompt patch as an editable diff. |