Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing long-text evaluation benchmarks, such as L-Eval and LongBench, focus on QA and summarization tasks. |
| Approach: | They propose a length-adaptable benchmark for evaluating the long-context understanding of large language models. |
| Outcome: | The proposed benchmarks do not cover ultralong settings (100k+ tokens) and are difficult to evaluate across different length ranges. |
Similar Papers
BAMBOO: A Comprehensive Benchmark for Evaluating Long Text Modeling Capacities of Large Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing long context models suffer from performance decline when the input text exceeds their length limit. |
| Approach: | They propose a multi-task long context benchmark to evaluate LLMs' long context ability using 10 datasets from 5 different NLP tasks. |
| Outcome: | The proposed model covers 5 domains and core capacities of large language models. |
TAIL: A Toolkit for Automatic and Realistic Long-Context Large Language Model Evaluation (2024.emnlp-demo)
Copied to clipboard
| Challenge: | Existing evaluation methods for long-context large language models are overly simplistic and require extensive human annotations. |
| Approach: | They propose an automatic toolkit to create realistic evaluation benchmarks . they use a document-grounded benchmark to generate question-answer pairs . |
| Outcome: | The proposed toolkit provides a way to create realistic evaluation benchmarks and visualize performance metrics of evaluated models. |
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for long-context capability are too synthetic and do not represent the real world usage of LLMs. |
| Approach: | They propose a length-controllable, real-life reflective benchmark that disentangles baseline knowledge from long-context capabilities. |
| Outcome: | Experiments show that the proposed benchmarks disentangle baseline knowledge from long-context capabilities. |
CLongEval: A Chinese Benchmark for Evaluating Long-Context Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Developing long-context LLMs with robust long-text capabilities is underdeveloped due to a lack of benchmarks. |
| Approach: | They propose a Chinese benchmark for evaluating long-context LLMs with Chinese capabilities. |
| Outcome: | The proposed model is based on 6 open-source LLMs and 2 commercial ones. |
KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches (2024.findings-emnlp)
Copied to clipboard
Jiayi Yuan, Hongyi Liu, Shaochen Zhong, Yu-Neng Chuang, Songchen Li, Guanchu Wang, Duy Le, Hongye Jin, Vipin Chaudhary, Zhaozhuo Xu, Zirui Liu, Xia Hu
| Challenge: | Long context capability is a crucial competency for large language models as it mitigates the human struggle to digest long-form texts. |
| Approach: | They propose to evaluate 10+ state-of-the-art approaches for long context-capable LLMs. |
| Outcome: | The proposed methods are compared against 10+ state-of-the-art approaches across seven categories of long context tasks. |
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding (2024.acl-long)
Copied to clipboard
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li
| Challenge: | Large language models (LLMs) can only handle texts a few thousand tokens long, limiting their applications on longer sequence inputs, such as books, reports, and codebases. |
| Approach: | They propose a bilingual, multi-task benchmark for long context understanding that extends context windows and more sophisticated memory mechanisms to improve models' long context capabilities. |
| Outcome: | The proposed model outperforms open-source models but struggles on longer contexts. |
LongTableBench: Benchmarking Long-Context Table Reasoning across Real-World Formats and Domains (2025.findings-emnlp)
Copied to clipboard
Liyao Li, Jiaming Tian, Hao Chen, Wentao Ye, Chao Ye, Haobo Wang, Ningtao Wang, Xing Fu, Gang Chen, Junbo Zhao
| Challenge: | Evaluating 52 LLMs reveals that only the strongest models maintain robust performance under increasing context lengths and format diversity. |
| Approach: | They propose a benchmark for evaluating long-context reasoning over semi-structured tables across diverse formats, tasks, and domains. |
| Outcome: | The proposed model outperforms compression-based approaches on tasks requiring semantic integration. |
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA (2024.emnlp-main)
Copied to clipboard
Minzheng Wang, Longze Chen, Fu Cheng, Shengyi Liao, Xinghua Zhang, Bingli Wu, Haiyang Yu, Nan Xu, Lei Zhang, Run Luo, Yunshui Li, Min Yang, Fei Huang, Yongbin Li
| Challenge: | Existing benchmarks for evaluating long-context language models employ irrelevant noise texts to artificially extend the length of test cases, diverging from the real-world scenarios of long-constituency applications. |
| Approach: | They propose a long-context benchmark, Loong, aligning with realistic scenarios through extended multi-document question answering (QA) . |
| Outcome: | The proposed model can scale up the context window of large language models to perform in-depth analysis of multiple long documents. |
LongSafety: Evaluating Long-Context Safety of Large Language Models (2025.acl-long)
Copied to clipboard
Yida Lu, Jiale Cheng, Zhexin Zhang, Shiyao Cui, Cunxiang Wang, Xiaotao Gu, Yuxiao Dong, Jie Tang, Hongning Wang, Minlie Huang
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable capabilities in understanding and generating long sequences. |
| Approach: | They propose a benchmark to evaluate LLM safety in open-ended long-context tasks . they find that relevant context and extended input sequences can exacerbate safety risks . |
| Outcome: | The proposed benchmark identifies significant safety vulnerabilities in 16 LLMs . strong safety performance in short-context scenarios does not correlate with safety in long-contact tasks . |
ELITR-Bench: A Meeting Assistant Benchmark for Long-Context Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Existing benchmarks for long-context LLMs focus on generic tasks that are not necessarily aligned with real-world applications. |
| Approach: | They propose to augment existing ELITR corpus by adding 271 manually crafted questions with their ground-truth answers and noisy versions of meeting transcripts altered to target different Word Error Rate levels. |
| Outcome: | The proposed benchmark augments the existing ELITR corpus by adding 271 manually crafted questions with ground-truth answers, as well as noisy versions of meeting transcripts altered to target different Word Error Rate levels. |