Papers with L-Eval

    1 papers
    Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks (2024.naacl-long)

    Copied to clipboard

    Challenge: Existing long-text evaluation benchmarks, such as L-Eval and LongBench, focus on QA and summarization tasks.
    Approach: They propose a length-adaptable benchmark for evaluating the long-context understanding of large language models.
    Outcome: The proposed benchmarks do not cover ultralong settings (100k+ tokens) and are difficult to evaluate across different length ranges.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations