Papers by Luca Gioacchini

    2 papers
    AutoPenBench: A Vulnerability Testing Benchmark for Generative Agents (2025.emnlp-industry)

    Copied to clipboard

    Challenge: LLM agents are promising for vulnerability testing, but lack benchmarks to evaluate and compare them.
    Approach: They propose an open-source benchmark for the evaluation of vulnerability testing agents that includes 33 tasks ranging from introductory exercises to actual vulnerable systems.
    Outcome: The proposed benchmark includes 33 tasks ranging from introductory exercises to actual vulnerable systems.
    AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents (2024.naacl-demo)

    Copied to clipboard

    Challenge: Existing benchmarks are narrow and simply compute overall task success.
    Approach: They propose a framework where both benchmarks and metrics are modular and easily extensible through well documented and easy-to-use APIs.
    Outcome: The proposed framework can track agent progress on two use cases and identify common failure points and refine the agent architecture to obtain a significant performance increase.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations