Papers by Ben Slater

    1 papers
    PredictaBoard: Benchmarking LLM Score Predictability (2025.findings-acl)

    Copied to clipboard

    Challenge: Large Language Models (LLMs) fail unpredictably, demonstrating inconsistent success in even basic common sense reasoning tasks.
    Approach: They propose a framework to evaluate the ability of score predictors to anticipate LLM errors on specific task instances from existing datasets.
    Outcome: The proposed framework evaluates the ability of score predictors to anticipate LLM errors on specific task instances from existing datasets.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations