Challenge: a new set of German-pretrained models are being released, but no established, diverse and systematic evaluation suite is available for them.
Approach: They assemble a Natural Language Understanding benchmark suite for the German language and evaluate 10 existing German-pretrained models.
Outcome: The proposed benchmark suite evaluates 10 German-pretrained models on 29 tasks . the results show that encoder models are good choices for most tasks, but not all .

Similar Papers

LLäMmlein: Transparent, Compact and Competitive German-Only Language Models from Scratch (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved remarkable success, yet this progress is predominantly centered on English.
Approach: They create two German-only decoder models from scratch and publish them for the (German) NLP research community to use.
Outcome: The two models performed competitively on the German SuperGLEBer benchmark, but performance improvements plateaued early during training, offering valuable insights into resource allocation for future models.
Superlim: A Swedish Language Understanding Evaluation Benchmark (2023.emnlp-main)

Copied to clipboard

Challenge: In this paper, we present a multi-task benchmark for Swedish language models . we address methodological challenges, such as mitigating the Anglocentric bias when creating datasets for a less-resourced language .
Approach: They propose a multi-task NLP benchmark for Swedish language models . they propose to use superlim to evaluate Swedish language model performance .
Outcome: The proposed benchmark does not approach ceiling performance on any of the tasks, suggesting it is difficult to implement.
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models (2025.naacl-long)

Copied to clipboard

Challenge: a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment.
Approach: They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation .
Outcome: The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks.
Navigating the Modern Evaluation Landscape: Considerations in Benchmarks and Frameworks for Large Language Models (LLMs) (2024.lrec-tutorials)

Copied to clipboard

Challenge: General-purpose Language Models have changed the world of Natural Language Processing, if not the world itself.
Approach: This tutorial will lay the foundations and explain the basics of evaluation and compare traditional methods to newly developed methods.
Outcome: The tutorial assumes little familiarity with metrics, datasets, prompts and benchmarks . it will compare traditional methods to newly developed methods .
German’s Next Language Model (2020.coling-main)

Copied to clipboard

Challenge: In this paper we compare the performance of our deep transformer based language models to existing models.
Approach: They present a set of BERT and ELECTRA based German language models, GBERT and GELECTRE.
Outcome: The proposed models outperform the previous best models on NER and classification tasks but are prohibitively large for many.
BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing multilingual benchmarks focus primarily on language understanding tasks.
Approach: They develop a multi-way multilingual benchmark that measures critical capabilities of large language models across languages.
Outcome: Extensive experiments on BenchMAX reveal uneven utilization of core capabilities across languages, emphasizing the performance gaps that scaling model size alone does not resolve.
RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark (2020.emnlp-main)

Copied to clipboard

Challenge: Modern scientific methodology is beginning to explore universal transformers as an independent object of study.
Approach: They propose a Russian general language understanding evaluation benchmark - Russian SuperGLUE . they provide a benchmark of nine tasks, human level evaluation and a leaderboard for the Russian language .
Outcome: The proposed benchmark provides nine tasks for the Russian language and human level evaluation and leaderboard of transformer models.
Evaluating Large Language Models with Enterprise Benchmarks (2025.naacl-industry)

Copied to clipboard

Challenge: Existing benchmarks lack domain-specific datasets for evaluating large language models . existing benchmarks often lack domain specific datasets, which can be difficult to convert to standardized metrics or regulatory issues.
Approach: They propose to use 25 publicly available domain-specific English benchmarks from diverse domains . they propose to combine a wide range of natural language processing tasks for holistic evaluation .
Outcome: The proposed framework includes 25 publicly available domain-specific English benchmarks from diverse enterprise domains like financial services, legal, climate, cyber security, and 2 public Japanese finance benchmarks.
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluations for Structured Knowledge (SK) understanding are non-rigorous and focus on a single type of SK.
Approach: They propose a structured knowledge understanding benchmark that includes four widely used structured knowledge forms.
Outcome: The proposed benchmark is based on four widely used structured knowledge forms . it includes a question, an answer, positive knowledge units, and noisy knowledge units .
ltzGLUE: Luxembourgish General Language Understanding Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: ltzGLUE is the first official NLU benchmark for Luxembourgish (LTZ) based on the popular GLUE benchmark for English.
Approach: They propose a new natural language understanding (NLU) benchmark for Luxembourgish based on the popular GLUE benchmark for English.
Outcome: The proposed model performs well across many languages and is based on the GLUE benchmark for English.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations