Challenge: Existing large language model (LLM) agents are not capable of performing research extension tasks autonomously.
Approach: They propose a benchmark to evaluate LLM agents' ability to extend existing AI research . they use extensions of 12 recently published research papers accompanied by domain expert-written instructions .
Outcome: The proposed benchmark evaluates 12 LLM agents implemented using aider and OpenHands.

Similar Papers

Agent Laboratory: Using LLM Agents as Research Assistants (2025.findings-emnlp)

Copied to clipboard

Challenge: Agent Laboratory is an autonomous LLM-based framework that can complete the entire research process.
Approach: Agent Laboratory is an autonomous LLM-based framework that can complete the entire research process.
Outcome: Agent Laboratory is an autonomous LLM-based framework that can complete the entire research process.
GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have been widely deployed as autonomous agents capable of following user instructions and making decisions in real-world applications.
Approach: They propose a benchmark to evaluate LLMs' ability to follow domain-oriented guidelines . they evaluate Lms on three critical aspects: adherence to diverse rules, robustness to rule updates .
Outcome: The proposed benchmark evaluates LLMs on three critical aspects: adherence to diverse rules, robustness to rule updates, and alignment with human preferences.
AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on single agentic capability, failing to capture long-horizon real-world scenarios.
Approach: They propose a benchmark that evaluates 6 agentic capabilities across 32 real-world scenarios.
Outcome: Experiments show that closed-source models outperform open-source model (48.4% vs 32.1%) integrating models with advanced scaffolds to form autonomous agents is a paradigm shift.
LMR-BENCH: Evaluating LLM Agent’s Ability on Reproducing Language Modeling Research (2025.emnlp-main)

Copied to clipboard

Challenge: Large language model (LLM) agents have demonstrated remarkable potential in advancing scientific discovery, but their capability in reproducing code from research papers remains underexplored.
Approach: They propose to evaluate LLM agents' ability to reproduce scientific research papers by analyzing code reproduction tasks from 23 research papers published in top-tier NLP venues.
Outcome: The proposed benchmark systematically evaluates the capability of large language model (LLM) agents on code reproduction from Language Modeling Research.
TinyScientist: An Interactive, Extensible, and Controllable Framework for Building Research Agents (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing research systems often design and use agentic workflows to perform research tasks such as ideation, scientific coding, review writing, and tree-based search.
Approach: They propose an open-source codebase, an interactive web demonstration, and a PyPI Python package to make state-of-the-art auto-research pipelines broadly accessible to every researcher and developer.
Outcome: The proposed framework adapts easily to new tools and supports iterative growth.
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories (2025.naacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated significant advances in computer science research . current agent-related applications include code writing, code base generation, code correction and more.
Approach: They propose a benchmark to assess the effectiveness of Large Language Models (LLMs) they propose GitHub agent framework that automates deployment of GitHub repositories .
Outcome: The proposed framework improves the deployment of GitHub code repositories and thereby boosts developer productivity.
LLM Agents Making Agent Tools (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) can perform multi-step tasks by dynamically utilising external software components.
Approach: They propose an agentic framework that autonomously transforms papers with code into LLM-compatible tools.
Outcome: The proposed framework outperforms current state-of-the-art software engineering agents in 80% of tasks and is openly available on GitHub.
Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments (2025.emnlp-main)

Copied to clipboard

Challenge: Enterprise systems are crucial for enhancing productivity and strategic growth, but data is fragmented across multiple sources and access controls are complex.
Approach: They propose a benchmark that simulates enterprise settings with 500 diverse tasks . they show that even the most capable models achieve only 41.8% task completion .
Outcome: The proposed benchmark shows that even the most capable models achieve only 41.8% task completion.
SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have made significant progress in writing code, but can they be used to reproduce results from research repositories?
Approach: They propose a benchmark to evaluate the capability of Large Language Models to reproduce results from research repositories.
Outcome: The benchmark aims to capture the realistic challenges faced by researchers working with machine learning and natural language processing repositories.
Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches lack flexibility to address diverse and ever-evolving user queries in open domains.
Approach: They propose to evaluate LLMs on open-domain knowledge that requires tools to solve diverse and ever-evolving user queries.
Outcome: The proposed system outperforms baselines in the open domain task-solving benchmark.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations