A Position Paper on the Automatic Generation of Machine Learning Leaderboards (2025.emnlp-main)
Copied to clipboard
| Challenge: | Automated leaderboard generation is a tool for comparing prior work with a tabular overview of experimental results. |
| Approach: | They propose an automatic leaderboard generation framework to standardise how the task is defined. |
| Outcome: | The proposed framework standardises how the ALG task is defined and proposes new directions . the proposed framework includes recommendations for datasets and metrics that promote fair evaluation . |
Similar Papers
Identification of Tasks, Datasets, Evaluation Metrics, and Numeric Scores for Scientific Leaderboards Construction (P19-1)
Copied to clipboard
| Challenge: | Recent years have witnessed a significant increase in laboratory-based evaluation benchmarks in many scientific disciplines. |
| Approach: | They propose to use NLP datasets to extract task, dataset, metric and score from NLP papers to build automatic leaderboards. |
| Outcome: | The proposed model outperforms baselines in the NLP domain by a large margin. |
LEGOBench: Scientific Leaderboard Generation Benchmark (2024.findings-emnlp)
Copied to clipboard
| Challenge: | a growing number of papers make it difficult to stay informed about the latest state-of-the-art research. |
| Approach: | They propose a benchmark to evaluate systems that generate scientific leaderboards . they use 22 years of submission data on arXiv and 11k machine learning leaderboard data on paperswithcode . |
| Outcome: | The proposed model shows significant performance gaps in the LEGOBench model . the model is based on a language model and four graph-based leaderboard generation task configuration . |
Efficient Performance Tracking: Leveraging Large Language Models for Automated Construction of Scientific Leaderboards (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing leaderboards are incomplete and some contain incorrect information. |
| Approach: | They propose a manually-curated Scientific Leaderboard dataset that overcomes these problems . they propose three experimental settings where TDM triples are fully defined, partially defined, or undefined . |
| Outcome: | The proposed system overcomes the shortcomings of existing leaderboard datasets . it can be used to evaluate and compare scientific methods, but it requires manual labor . |
Bidimensional Leaderboards: Generate and Evaluate Language Hand in Hand (2022.naacl-main)
Copied to clipboard
Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan, Jacob Morrison, Alexander Fabbri, Yejin Choi, Noah A. Smith
| Challenge: | Recent advances on models and metrics should benefit and inform each other, authors argue . bidimensional leaderboards allow for fast, accurate evaluation of language generation models . |
| Approach: | They propose a bidimensional leaderboard that tracks progress in language generation models and metrics for their evaluation. |
| Outcome: | The proposed leaderboards track progress in language generation models and metrics for their evaluation. |
ExplainaBoard: An Explainable Leaderboard for NLP (2021.acl-demo)
Copied to clipboard
Pengfei Liu, Jinlan Fu, Yang Xiao, Weizhe Yuan, Shuaichen Chang, Junqi Dai, Yixin Liu, Zihuiwen Ye, Graham Neubig
| Challenge: | Using leaderboards, researchers can track the performance of various systems on various NLP tasks. |
| Approach: | They propose a new conceptualization and implementation of NLP evaluation using a leaderboard. |
| Outcome: | The ExplainaBoard is an evaluation tool for natural language processing (NLP) it covers more than 400 systems, 50 datasets, 40 languages, and 12 tasks. |
Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards? (2021.acl-long)
Copied to clipboard
| Challenge: | Rather than replacing leaderboards, we advocate a re-imagining of the model to highlight if and where progress is made. |
| Approach: | They propose a Bayesian leaderboard model where latent subject skill and latent item difficulty predict correct responses. |
| Outcome: | The proposed model can guide what to annotate, identify annotation errors, detect overfitting, and identify informative examples. |
Striking Gold in Advertising: Standardization and Exploration of Ad Text Generation (2024.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks and problem sets for automatic ad text generation are lacking . however, the growing volume of search queries has fueled research on the automatic generation of ads. |
| Approach: | They propose to standardize the task of automatic ad text generation (ATG) using a benchmark dataset, CAMERA, to enable the utilization of multi-modal information and facilitate industry-wise evaluations. |
| Outcome: | The proposed dataset standardizes the task of automatic ad text generation (ATG) it shows that existing metrics align with human evaluations and that the proposed methods can be used to improve the quality of the results. |
Academics Can Contribute to Domain-Specialized Language Models (2024.emnlp-main)
Copied to clipboard
Mark Dredze, Genta Winata, Prabhanjan Kambadur, Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, David Rosenberg, Sebastian Gehrmann
| Challenge: | Commercially available models dominate academic leaderboards, focusing on creating and adapting general-purpose models . however, general- purpose models often underperform in specialized domains, and domain-specific models yield superior results. |
| Approach: | They advocate for a renewed focus on developing and evaluating domain- and task-specific models . they advocate for an adapted or adapted model that can be used to improve academic leaderboard standings . |
| Outcome: | The proposed model can do well on professional and linguistic examinations, college-level knowledge questions, and collections of reasoning tasks. |
Targeting the Benchmark: On Methodology in Current Natural Language Processing Research (2021.acl-short)
Copied to clipboard
| Challenge: | a language benchmark is a task devised that is restricted enough to be managable with current methods, but is deemed challenging enough to serve as a benchmark. |
| Approach: | They propose to use a language task as a benchmark and a baseline model to argue it is challenging enough to be a good one. |
| Outcome: | The proposed language benchmarks are based on a dataset and a language task . the proposed benchmarks can be used to measure progress towards the goal of the research . |
Unveiling the Art of Heading Design: A Harmonious Blend of Summarization, Neology, and Algorithm (2024.findings-acl)
Copied to clipboard
| Challenge: | Creating an appealing heading is crucial for attracting readers and marketing work or products. |
| Approach: | They propose a benchmark to measure the quality of heading generation using summarization, neology, and algorithm metrics. |
| Outcome: | The proposed benchmark compared 6,653 abstracts with corresponding descriptions and acronyms and found that it excels across summarization, neology, and algorithm aspects. |