| Challenge: | Natural Language Understanding (NLU) benchmarks are costly to develop and language-dependent . basqueGLUE is the first benchmark for Basque, a less-resourced language . |
| Approach: | They propose a benchmark for Basque, a less-resourced language, using existing datasets. |
| Outcome: | The proposed benchmarks take into account a wide and diverse set of NLU tasks that require some form of language understanding beyond the detection of superficial clues. |
Similar Papers
JGLUE: Japanese General Language Understanding Evaluation (2022.lrec-1)
Copied to clipboard
| Challenge: | There is no benchmark for Japanese to evaluate and analyze NLU ability from different perspectives. |
| Approach: | They build a Japanese NLU benchmark from scratch without translation to measure general NLU ability in Japanese. |
| Outcome: | a Japanese NLU benchmark is built from scratch without translation to measure general NLU ability in Japanese. |
Superlim: A Swedish Language Understanding Evaluation Benchmark (2023.emnlp-main)
Copied to clipboard
Aleksandrs Berdicevskis, Gerlof Bouma, Robin Kurtz, Felix Morger, Joey Öhman, Yvonne Adesam, Lars Borin, Dana Dannélls, Markus Forsberg, Tim Isbister, Anna Lindahl, Martin Malmsten, Faton Rekathati, Magnus Sahlgren, Elena Volodina, Love Börjeson, Simon Hengchen, Nina Tahmasebi
| Challenge: | In this paper, we present a multi-task benchmark for Swedish language models . we address methodological challenges, such as mitigating the Anglocentric bias when creating datasets for a less-resourced language . |
| Approach: | They propose a multi-task NLP benchmark for Swedish language models . they propose to use superlim to evaluate Swedish language model performance . |
| Outcome: | The proposed benchmark does not approach ceiling performance on any of the tasks, suggesting it is difficult to implement. |
ltzGLUE: Luxembourgish General Language Understanding Evaluation (2026.findings-acl)
Copied to clipboard
Alistair Plum, Felicia Körner, Anne-Marie Lutgen, Laura Bernardy, Fred Philippy, Emilia Milano, Nils Rehlinger, Cedric Lothritz, Tharindu Ranasinghe, Barbara Plank, Christoph Purschke
| Challenge: | ltzGLUE is the first official NLU benchmark for Luxembourgish (LTZ) based on the popular GLUE benchmark for English. |
| Approach: | They propose a new natural language understanding (NLU) benchmark for Luxembourgish based on the popular GLUE benchmark for English. |
| Outcome: | The proposed model performs well across many languages and is based on the GLUE benchmark for English. |
UINAUIL: A Unified Benchmark for Italian Natural Language Understanding (2023.acl-demo)
Copied to clipboard
| Challenge: | a benchmark of six tasks for Italian Natural Language Understanding is presented . large language models (LLMs) have revolutionized the field of natural language processing . a few benchmarks exist for non-English languages, but only a handful are available for nonEnglish languages . |
| Approach: | They introduce a benchmark for Italian Natural Language Understanding that harmonizes the data format and exposes functionalities to facilitate data manipulation and evaluation of custom models. |
| Outcome: | The proposed benchmarks are based on the European Language Grid and available models in Italian and multilingual languages. |
XNLIeu: a dataset for cross-lingual NLI in Basque (2024.naacl-long)
Copied to clipboard
| Challenge: | XNLI is a popular benchmark used to evaluate cross-lingual Natural Language Understanding (NLU) in languages such as English, Basque and other low-resource languages. |
| Approach: | They expand XNLI to include Basque, a low-resource language that can benefit from transfer-learning approaches. |
| Outcome: | The proposed dataset includes Basque, a low-resource language that can benefit from transfer-learning approaches. |
skLEP: A Slovak General Language Understanding Benchmark (2025.findings-acl)
Copied to clipboard
Marek Suppa, Andrej Ridzik, Daniel Hládek, Tomáš Javůrek, Viktória Ondrejová, Kristína Sásiková, Martin Tamajka, Marian Simko
| Challenge: | skLEP is the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding models. |
| Approach: | They introduce a benchmark specifically designed for evaluating Slovak natural language understanding models. |
| Outcome: | The proposed benchmark covers nine tasks that span token-level, sentence-pair, document-level tasks. |
KLEJ: Comprehensive Benchmark for Polish Language Understanding (2020.acl-main)
Copied to clipboard
| Challenge: | Recent introduction of robust, general-purpose models for fine-tuning has enabled improvements in general natural language understanding (NLU) but such benchmarks are only available for a handful of languages. |
| Approach: | They propose a multi-task benchmark for the Polish language understanding with an online leaderboard . they also propose GLUE, a task for named entity recognition and sentiment analysis . |
| Outcome: | The proposed model performs best on three out of nine tasks in the Polish language . the proposed model is also used in an e-commerce domain to analyze the sentiments of users . |
CLUE: A Chinese Language Understanding Evaluation Benchmark (2020.coling-main)
Copied to clipboard
Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, Zhenzhong Lan
| Challenge: | Existing language evaluation benchmarks for English are limited to English . lack of such benchmarks makes it difficult to replicate success in other languages . |
| Approach: | They introduce a large-scale Chinese language understanding evaluation benchmark . the benchmark uses a set of current state-of-the-art pre-trained Chinese models . |
| Outcome: | The first large-scale Chinese Language Understanding Evaluation (CLUE) benchmark is released . the benchmark evaluates models across a wide range of tasks on original Chinese text . existing language evaluation benchmarks are mostly limited to English . |
NLPre: A Revised Approach towards Language-centric Benchmarking of Natural Language Preprocessing Systems (2024.lrec-main)
Copied to clipboard
| Challenge: | GLUE benchmarking system enables ongoing evaluation of multiple NLPre tools while credibly tracking their performance. |
| Approach: | They propose a language-centric benchmarking system that enables ongoing evaluation of multiple NLPre tools while credibly tracking their performance. |
| Outcome: | The proposed system is configured for Polish and integrated with the thoroughly assembled NLPre-PL benchmark. |
RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark (2020.emnlp-main)
Copied to clipboard
Tatiana Shavrina, Alena Fenogenova, Emelyanov Anton, Denis Shevelev, Ekaterina Artemova, Valentin Malykh, Vladislav Mikhailov, Maria Tikhonova, Andrey Chertok, Andrey Evlampiev
| Challenge: | Modern scientific methodology is beginning to explore universal transformers as an independent object of study. |
| Approach: | They propose a Russian general language understanding evaluation benchmark - Russian SuperGLUE . they provide a benchmark of nine tasks, human level evaluation and a leaderboard for the Russian language . |
| Outcome: | The proposed benchmark provides nine tasks for the Russian language and human level evaluation and leaderboard of transformer models. |