Papers by Nicholas Edwards

    1 papers
    RExBench: Can coding agents autonomously implement AI research extensions? (2026.acl-long)

    Copied to clipboard

    Challenge: Existing large language model (LLM) agents are not capable of performing research extension tasks autonomously.
    Approach: They propose a benchmark to evaluate LLM agents' ability to extend existing AI research . they use extensions of 12 recently published research papers accompanied by domain expert-written instructions .
    Outcome: The proposed benchmark evaluates 12 LLM agents implemented using aider and OpenHands.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations