| Challenge: | GLUE, SuperGLUE and RussianSuperGLUE benchmarks are arbitrary sets of tasks that are not generalized. |
| Approach: | They propose a theoretical instrument and an algorithm to calculate similarity between benchmark tasks . they use relative performance of the "students" on a given task to determine similarity . |
| Outcome: | The proposed model reduces the number of evaluation tasks while maintaining high validation quality. |
Similar Papers
RPD: A Distance Function Between Word Embeddings (2020.acl-srw)
Copied to clipboard
| Challenge: | Existing word embeddings are poorly understood, but little is known about how they differ between different sets of word embeds. |
| Approach: | They propose a metric called Relative Pairwise Inner Product Distance to quantify the distance between different word embeddings. |
| Outcome: | The proposed metric measures the distance between different sets of embeddings and investigates the influence of different training processes and corpora. |
MUSTS: MUltilingual Semantic Textual Similarity Benchmark (2025.acl-short)
Copied to clipboard
| Challenge: | Existing benchmarks for semantic textual similarity (STS) are limited to high-resource languages and do not include datasets annotated focusing on relatedness instead of similarity. |
| Approach: | They propose to evaluate multilingual semantic textual similarity benchmarks which span 13 languages and annotated datasets to evaluate and compare them. |
| Outcome: | The proposed method is the most comprehensive benchmark of multilingual STS methods. |
Cross-Task Generalization Abilities of Large Language Models (2024.naacl-srw)
Copied to clipboard
| Challenge: | a thesis proposal advocates for the crucial role of cross-task generalization in NLP systems. |
| Approach: | They propose to benchmark cross-task generalization abilities with diverse NLP tasks . they also propose to develop model architectures for improving cross- task generalization . |
| Outcome: | This paper compares cross-task generalization abilities with diverse NLP tasks . it also analyzes and predicts the generalization landscape of current state-of-the-art large language models . |
MATCH: Task-Driven Code Evaluation through Contrastive Learning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | GitHub Copilot generates 46% of the code on GitHub. |
| Approach: | They propose a reference-free metric that uses Contrastive Learning to generate meaningful embeddings for code and natural language task descriptions. |
| Outcome: | This paper compares the performance of a new similarity score with existing metrics. |
Towards Explainable Evaluation of Language Models on the Semantic Similarity of Visual Concepts (2022.coling-1)
Copied to clipboard
Maria Lymperaiou, George Manoliadis, Orfeas Menis Mastromichalakis, Edmund G. Dervakos, Giorgos Stamou
| Challenge: | Recent advances in NLP research have focused on robustness and explainability issues of their evaluation strategies. |
| Approach: | They propose to use pre-trained transformers to evaluate semantic similarity for visual vocabularies . they propose to provide explainable metrics for understanding the quality of retrieved instances . |
| Outcome: | The proposed metrics highlight inabilities of widely used evaluation methods and highlight weaknesses in learned linguistic representations. |
In Benchmarks We Trust ... Or Not? (2025.emnlp-main)
Copied to clipboard
Ine Gevers, Victor De Marez, Jens Van Nooten, Jens Lemmens, Andriy Kosar, Ehsan Lotfi, Nikolay Banar, Pieter Fivez, Luna De Bruyne, Walter Daelemans
| Challenge: | Existing benchmarks for Large Language Models (LLMs) are inadequate and lack a clear solution. |
| Approach: | They propose checklists to cover all aspects of benchmarking issues, both for benchmark creation and usage. |
| Outcome: | The proposed checklists cover all aspects of benchmarking issues, both for benchmark creation and usage. |
SuperTweetEval: A Challenging, Unified and Heterogeneous Benchmark for Social Media NLP Research (2023.findings-emnlp)
Copied to clipboard
Dimosthenis Antypas, Asahi Ushio, Francesco Barbieri, Leonardo Neves, Kiamehr Rezaee, Luis Espinosa-Anke, Jiaxin Pei, Jose Camacho-Collados
| Challenge: | specialised language models (LMs) have shown to exhibit lower perplexity and higher downstream performance across the board. |
| Approach: | They propose a benchmark for NLP evaluation in social media, SuperTweetEval. |
| Outcome: | The proposed benchmark shows that social media models perform better when compared to general-purpose models, metrics and benchmarks. |
A large-scale computational study of content preservation measures for text style transfer and paraphrase generation (2022.acl-srw)
Copied to clipboard
| Challenge: | Text style transfer and paraphrases generation are growing areas of NLP . many researchers still use BLEU-like measures to evaluate content preservation . |
| Approach: | They compare 57 different measures based on different principles on 19 annotated datasets . they find that measures relying on cross-encoder models outperform alternative approaches . |
| Outcome: | The proposed methods outperform traditional methods on 19 datasets. |
SMATCH++: Standardized and Extended Evaluation of Semantic Graphs (2023.findings-eacl)
Copied to clipboard
| Challenge: | Existing graph-alignment metrics that measure graph distances are not reliable, we show . metric is spread out and does not provide upper bounds for extended tasks. |
| Approach: | They propose a metric to measure a distance between graphs by aligning nodes and counting matching graph triples. |
| Outcome: | The proposed method reduces search space and improves scoring by reducing the number of errors. |
WYWEB: A NLP Evaluation Benchmark For Classical Chinese (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for classical Chinese are inadequate to evaluate performance of different NLP models. |
| Approach: | They propose an evaluation benchmark for classical Chinese NLP, which evaluates existing models. |
| Outcome: | The proposed benchmark evaluates the performance of existing models in classical Chinese. |