Vygotsky Distance: Measure for Benchmark Task Similarity (2024.lrec-main)

Copied to clipboard

Challenge: GLUE, SuperGLUE and RussianSuperGLUE benchmarks are arbitrary sets of tasks that are not generalized.
Approach: They propose a theoretical instrument and an algorithm to calculate similarity between benchmark tasks . they use relative performance of the "students" on a given task to determine similarity .
Outcome: The proposed model reduces the number of evaluation tasks while maintaining high validation quality.

Similar Papers

RPD: A Distance Function Between Word Embeddings (2020.acl-srw)

Copied to clipboard

Challenge: Existing word embeddings are poorly understood, but little is known about how they differ between different sets of word embeds.
Approach: They propose a metric called Relative Pairwise Inner Product Distance to quantify the distance between different word embeddings.
Outcome: The proposed metric measures the distance between different sets of embeddings and investigates the influence of different training processes and corpora.
MUSTS: MUltilingual Semantic Textual Similarity Benchmark (2025.acl-short)

Copied to clipboard

Challenge: Existing benchmarks for semantic textual similarity (STS) are limited to high-resource languages and do not include datasets annotated focusing on relatedness instead of similarity.
Approach: They propose to evaluate multilingual semantic textual similarity benchmarks which span 13 languages and annotated datasets to evaluate and compare them.
Outcome: The proposed method is the most comprehensive benchmark of multilingual STS methods.
Cross-Task Generalization Abilities of Large Language Models (2024.naacl-srw)

Copied to clipboard

Challenge: a thesis proposal advocates for the crucial role of cross-task generalization in NLP systems.
Approach: They propose to benchmark cross-task generalization abilities with diverse NLP tasks . they also propose to develop model architectures for improving cross- task generalization .
Outcome: This paper compares cross-task generalization abilities with diverse NLP tasks . it also analyzes and predicts the generalization landscape of current state-of-the-art large language models .
MATCH: Task-Driven Code Evaluation through Contrastive Learning (2025.findings-emnlp)

Copied to clipboard

Challenge: GitHub Copilot generates 46% of the code on GitHub.
Approach: They propose a reference-free metric that uses Contrastive Learning to generate meaningful embeddings for code and natural language task descriptions.
Outcome: This paper compares the performance of a new similarity score with existing metrics.
Towards Explainable Evaluation of Language Models on the Semantic Similarity of Visual Concepts (2022.coling-1)

Copied to clipboard

Challenge: Recent advances in NLP research have focused on robustness and explainability issues of their evaluation strategies.
Approach: They propose to use pre-trained transformers to evaluate semantic similarity for visual vocabularies . they propose to provide explainable metrics for understanding the quality of retrieved instances .
Outcome: The proposed metrics highlight inabilities of widely used evaluation methods and highlight weaknesses in learned linguistic representations.
In Benchmarks We Trust ... Or Not? (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for Large Language Models (LLMs) are inadequate and lack a clear solution.
Approach: They propose checklists to cover all aspects of benchmarking issues, both for benchmark creation and usage.
Outcome: The proposed checklists cover all aspects of benchmarking issues, both for benchmark creation and usage.
SuperTweetEval: A Challenging, Unified and Heterogeneous Benchmark for Social Media NLP Research (2023.findings-emnlp)

Copied to clipboard

Challenge: specialised language models (LMs) have shown to exhibit lower perplexity and higher downstream performance across the board.
Approach: They propose a benchmark for NLP evaluation in social media, SuperTweetEval.
Outcome: The proposed benchmark shows that social media models perform better when compared to general-purpose models, metrics and benchmarks.
A large-scale computational study of content preservation measures for text style transfer and paraphrase generation (2022.acl-srw)

Copied to clipboard

Challenge: Text style transfer and paraphrases generation are growing areas of NLP . many researchers still use BLEU-like measures to evaluate content preservation .
Approach: They compare 57 different measures based on different principles on 19 annotated datasets . they find that measures relying on cross-encoder models outperform alternative approaches .
Outcome: The proposed methods outperform traditional methods on 19 datasets.
SMATCH++: Standardized and Extended Evaluation of Semantic Graphs (2023.findings-eacl)

Copied to clipboard

Challenge: Existing graph-alignment metrics that measure graph distances are not reliable, we show . metric is spread out and does not provide upper bounds for extended tasks.
Approach: They propose a metric to measure a distance between graphs by aligning nodes and counting matching graph triples.
Outcome: The proposed method reduces search space and improves scoring by reducing the number of errors.
WYWEB: A NLP Evaluation Benchmark For Classical Chinese (2023.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for classical Chinese are inadequate to evaluate performance of different NLP models.
Approach: They propose an evaluation benchmark for classical Chinese NLP, which evaluates existing models.
Outcome: The proposed benchmark evaluates the performance of existing models in classical Chinese.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations