Challenge: Existing leaderboards are incomplete and some contain incorrect information.
Approach: They propose a manually-curated Scientific Leaderboard dataset that overcomes these problems . they propose three experimental settings where TDM triples are fully defined, partially defined, or undefined .
Outcome: The proposed system overcomes the shortcomings of existing leaderboard datasets . it can be used to evaluate and compare scientific methods, but it requires manual labor .

Similar Papers

Identification of Tasks, Datasets, Evaluation Metrics, and Numeric Scores for Scientific Leaderboards Construction (P19-1)

Copied to clipboard

Challenge: Recent years have witnessed a significant increase in laboratory-based evaluation benchmarks in many scientific disciplines.
Approach: They propose to use NLP datasets to extract task, dataset, metric and score from NLP papers to build automatic leaderboards.
Outcome: The proposed model outperforms baselines in the NLP domain by a large margin.
LEGOBench: Scientific Leaderboard Generation Benchmark (2024.findings-emnlp)

Copied to clipboard

Challenge: a growing number of papers make it difficult to stay informed about the latest state-of-the-art research.
Approach: They propose a benchmark to evaluate systems that generate scientific leaderboards . they use 22 years of submission data on arXiv and 11k machine learning leaderboard data on paperswithcode .
Outcome: The proposed model shows significant performance gaps in the LEGOBench model . the model is based on a language model and four graph-based leaderboard generation task configuration .
A Position Paper on the Automatic Generation of Machine Learning Leaderboards (2025.emnlp-main)

Copied to clipboard

Challenge: Automated leaderboard generation is a tool for comparing prior work with a tabular overview of experimental results.
Approach: They propose an automatic leaderboard generation framework to standardise how the task is defined.
Outcome: The proposed framework standardises how the ALG task is defined and proposes new directions . the proposed framework includes recommendations for datasets and metrics that promote fair evaluation .
MetaLead: A Comprehensive Human-Curated Leaderboard Dataset for Transparent Reporting of Machine Learning Experiments (2026.eacl-long)

Copied to clipboard

Challenge: Existing leaderboards capture only the best results from each paper and have limited metadata.
Approach: They propose to create a fully human-annotated ML Leaderboard dataset that captures all experimental results and contains extra metadata.
Outcome: The MetaLead dataset captures all experimental results and contains extra metadata for cross-domain evaluation.
ExplainaBoard: An Explainable Leaderboard for NLP (2021.acl-demo)

Copied to clipboard

Challenge: Using leaderboards, researchers can track the performance of various systems on various NLP tasks.
Approach: They propose a new conceptualization and implementation of NLP evaluation using a leaderboard.
Outcome: The ExplainaBoard is an evaluation tool for natural language processing (NLP) it covers more than 400 systems, 50 datasets, 40 languages, and 12 tasks.
ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition (2026.findings-acl)

Copied to clipboard

Challenge: Large language models have shown potential in assisting scientific research, yet their ability to discover high-quality research hypotheses remains unexamined due to the lack of a dedicated benchmark.
Approach: They propose a benchmark for evaluating large language models on a sufficient set of scientific discovery sub-tasks.
Outcome: The proposed framework extracts critical components from papers across 12 disciplines with expert validation confirming its accuracy.
Academics Can Contribute to Domain-Specialized Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Commercially available models dominate academic leaderboards, focusing on creating and adapting general-purpose models . however, general- purpose models often underperform in specialized domains, and domain-specific models yield superior results.
Approach: They advocate for a renewed focus on developing and evaluating domain- and task-specific models . they advocate for an adapted or adapted model that can be used to improve academic leaderboard standings .
Outcome: The proposed model can do well on professional and linguistic examinations, college-level knowledge questions, and collections of reasoning tasks.
SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluations of large language models fail to reflect fine-grained capabilities . existing benchmarks are manually curated or domain-generic, limiting scalability and alignment with real use cases.
Approach: They propose a framework that allows custom construction of benchmarks from large-scale scientific data to evaluate application-specific scientific capabilities in LLMs.
Outcome: The proposed framework reveals fine-grained differences in scientific capabilities that standard benchmarks overlook . it allows custom construction of benchmarks from large-scale scientific data to evaluate application-specific capabilities in LLMs.
Bidimensional Leaderboards: Generate and Evaluate Language Hand in Hand (2022.naacl-main)

Copied to clipboard

Challenge: Recent advances on models and metrics should benefit and inform each other, authors argue . bidimensional leaderboards allow for fast, accurate evaluation of language generation models .
Approach: They propose a bidimensional leaderboard that tracks progress in language generation models and metrics for their evaluation.
Outcome: The proposed leaderboards track progress in language generation models and metrics for their evaluation.
ResearchArena: Benchmarking Large Language Models’ Ability to Collect and Organize Information as Research Agents (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models excel across many natural language processing tasks but face challenges in domain-specific, analytical tasks such as conducting research surveys.
Approach: They propose a benchmark to evaluate LLMs' capabilities in conducting research surveys.
Outcome: The proposed benchmark is designed to evaluate LLMs' capabilities in conducting research surveys.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations