RusConText Benchmark: A Russian Language Evaluation Benchmark for Understanding Context (2025.acl-srw)
Copied to clipboard
| Challenge: | a new context understanding benchmark is proposed for short-context understanding in Russian . the benchmarks focus on broad reasoning tasks or long-concept comprehension, but are limited in their ability to perceive subtle nuances of context. |
| Approach: | They propose a new benchmark for evaluating short-context understanding in Russian . they propose to use four tasks to assess model performance from a specific perspective . |
| Outcome: | The proposed benchmark is adapted to Russian-language data. |
Similar Papers
RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark (2020.emnlp-main)
Copied to clipboard
Tatiana Shavrina, Alena Fenogenova, Emelyanov Anton, Denis Shevelev, Ekaterina Artemova, Valentin Malykh, Vladislav Mikhailov, Maria Tikhonova, Andrey Chertok, Andrey Evlampiev
| Challenge: | Modern scientific methodology is beginning to explore universal transformers as an independent object of study. |
| Approach: | They propose a Russian general language understanding evaluation benchmark - Russian SuperGLUE . they provide a benchmark of nine tasks, human level evaluation and a leaderboard for the Russian language . |
| Outcome: | The proposed benchmark provides nine tasks for the Russian language and human level evaluation and leaderboard of transformer models. |
Can Large Language Models Understand Context? (2024.findings-eacl)
Copied to clipboard
Yilun Zhu, Joel Moniz, Shruti Bhargava, Jiarui Lu, Dhivya Piraviperumal, Site Li, Yuan Zhang, Hong Yu, Bo-Hsiang Tseng
| Challenge: | Existing evaluation methodologies for Large Language Models (LLMs) have been inadequate to evaluate their ability to understand contextual features. |
| Approach: | They propose a benchmark to assess large language models' ability to understand context by adapting existing datasets to suit their evaluation. |
| Outcome: | The proposed model performs better under the in-context learning pretraining scenario than state-of-the-art models. |
ePiC: Employing Proverbs in Context as a Benchmark for Abstract Language Understanding (2022.acl-long)
Copied to clipboard
| Challenge: | Large language models have shown exciting progress on several NLP benchmarks . however, evaluating their ability for complex analogical reasoning remains under-explored . |
| Approach: | They propose a dataset of narratives for employing proverbs in context as a benchmark for abstract language understanding. |
| Outcome: | The proposed dataset provides fine-grained annotation of aligned spans between proverbs and narratives and contains minimal overlaps between narratives with proverb . the results show that large language models struggle on these tasks compared to humans, and these tasks pose multiple learning challenges. |
CoDA21: Evaluating Language Understanding Capabilities of NLP Models With Context-Definition Alignment (2022.acl-short)
Copied to clipboard
| Challenge: | Pretrained language models (PLMs) have achieved superhuman performance on many benchmarks, creating a need for harder tasks. |
| Approach: | They propose a benchmark that measures natural language understanding (NLU) abilities of pretrained language models. |
| Outcome: | The proposed benchmark measures the ability of pretrained language models to perform on many tasks. |
NeedleChain: Measuring Intact Context Comprehension Capability of Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for context understanding embed query-irrelevant content . this shifts evaluation toward retrieving relevant snippets rather than fully integrating all provided information. |
| Approach: | They propose a benchmark to evaluate whether models can faithfully incorporate all given evidence . they propose 'needlechain' benchmark to test whether models incorporate all available information . |
| Outcome: | The proposed benchmarks overestimate the ability of large language models to integrate all given evidence when the context is entirely query-relevant. |
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding (2024.acl-long)
Copied to clipboard
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li
| Challenge: | Large language models (LLMs) can only handle texts a few thousand tokens long, limiting their applications on longer sequence inputs, such as books, reports, and codebases. |
| Approach: | They propose a bilingual, multi-task benchmark for long context understanding that extends context windows and more sophisticated memory mechanisms to improve models' long context capabilities. |
| Outcome: | The proposed model outperforms open-source models but struggles on longer contexts. |
Marathon: A Race Through the Realm of Long Context with Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Existing long-context benchmarks do not accurately evaluate large language models’ comprehension and reasoning abilities in extended texts. |
| Approach: | They propose a new evaluation benchmark that adopts a multiple-choice question format and uses a multi-choke question format to assess the comprehension and reasoning skills of large language models. |
| Outcome: | The proposed benchmark provides a rapid, precise, and unbiased appraisal of the long-context comprehension skills of large language models. |
BeDiscovER: The Benchmark of Discourse Understanding in the Era of Reasoning Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | BeDiscovER evaluates the discourse-level knowledge of modern LLMs . state-of-the-art models exhibit strong performance in arithmetic aspect of temporal reasoning, but struggle with long-dependency reasoning and some subtle semantic and discourse phenomena, such as rhetorical relation classification. |
| Approach: | They evaluate open-source LLMs Qwen3 series, DeepSeek-R1, and frontier reasoning model GPT-5-mini on BeDiscovER . they find that models exhibit strong performance in arithmetic aspect of temporal reasoning, but struggle with long-dependency reasoning and some subtle semantic and discourse phenomena . |
| Outcome: | The proposed framework evaluates open-source LLMs Qwen3 series, DeepSeek-R1, and frontier reasoning model GPT-5-mini. |
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing multilingual benchmarks focus primarily on language understanding tasks. |
| Approach: | They develop a multi-way multilingual benchmark that measures critical capabilities of large language models across languages. |
| Outcome: | Extensive experiments on BenchMAX reveal uneven utilization of core capabilities across languages, emphasizing the performance gaps that scaling model size alone does not resolve. |
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)
Copied to clipboard
Bolei Ma, Yuting Li, Wei Zhou, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, Barbara Plank
| Challenge: | linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions. |
| Approach: | They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications. |
| Outcome: | The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models . |