| Challenge: | Modern language models exhibit powerful but brittle behavior, leading to larger and more diverse benchmarks. |
| Approach: | They propose to use anchor points to select small subsets of a language model-prompt dataset to capture model behavior across the entire dataset. |
| Outcome: | The proposed technique outperforms standard benchmarks in language models with 1-30 anchor points . the proposed technique can be used to compare models on different regions of the dataset . |
Similar Papers
In Benchmarks We Trust ... Or Not? (2025.emnlp-main)
Copied to clipboard
Ine Gevers, Victor De Marez, Jens Van Nooten, Jens Lemmens, Andriy Kosar, Ehsan Lotfi, Nikolay Banar, Pieter Fivez, Luna De Bruyne, Walter Daelemans
| Challenge: | Existing benchmarks for Large Language Models (LLMs) are inadequate and lack a clear solution. |
| Approach: | They propose checklists to cover all aspects of benchmarking issues, both for benchmark creation and usage. |
| Outcome: | The proposed checklists cover all aspects of benchmarking issues, both for benchmark creation and usage. |
Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Language models (LMs) often rely on spurious correlations rather than causally relevant features to improve accuracy and generalizability. |
| Approach: | They propose a benchmark that categorizes shortcuts into occurrence, style, and concept . they aim to explore the nuanced ways shortcuts influence the performance of LMs . |
| Outcome: | The proposed benchmark categorizes shortcuts into occurrence, style, and concept . it systematically investigates models’ resilience and susceptibilities to sophisticated shortcuts . |
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards (2024.acl-long)
Copied to clipboard
Norah Alzahrani, Hisham Alyahya, Yazeed Alnumay, Sultan AlRashed, Shaykhah Alsubaie, Yousef Almushayqih, Faisal Mirza, Nouf Alotaibi, Nora Al-Twairesh, Areeb Alowisheq, M Saiful Bari, Haidar Khan
| Challenge: | Existing leaderboards are often taken at face value, but this is costly . a recent study shows that minor perturbations to the benchmark result in rankings up to 8 positions. |
| Approach: | They propose to use a *hybrid* scoring method for answer selection for large language models . they find that minor perturbations to the benchmark result in rankings changes . |
| Outcome: | The proposed model is a hybrid scoring method, the authors argue . the proposed model could be used to improve the performance of large language models . |
Small Language Models Are Good Too: An Empirical Study of Zero-Shot Classification (2024.lrec-main)
Copied to clipboard
| Challenge: | Using small language models, we challenge the dominance of large models in text classification by prompting. |
| Approach: | They compare the performance of small and large language models in a zero-shot context using different architectures and scoring functions. |
| Outcome: | The proposed model outperforms large models in a zero-shot context. |
Beyond the Tip of the Iceberg: Assessing Coherence of Text Classifiers (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Large-scale, pre-trained language models achieve human-level and superhuman accuracy on existing language understanding tasks, but statistical bias in benchmark data and probing studies has recently called into question their true capabilities. |
| Approach: | They propose to evaluate systems through a measure of prediction coherence by using two existing language understanding benchmarks with different properties to demonstrate its versatility. |
| Outcome: | The proposed evaluation framework is quick, effective, and versatile to provide insight into the coherence of machines’ predictions. |
Chasing Random: Instruction Selection Strategies Fail to Generalize (2025.findings-naacl)
Copied to clipboard
| Challenge: | Prior work has shown that language models can be tuned to follow user instructions using only a small set of high-quality instructions. |
| Approach: | They analyze popular selection strategies across different datasets and benchmarks to find out whether they generalize poorly. |
| Outcome: | The proposed methods outperform random baselines and cost-performance trade-offs on the full dataset and a random subset. |
Navigating the Modern Evaluation Landscape: Considerations in Benchmarks and Frameworks for Large Language Models (LLMs) (2024.lrec-tutorials)
Copied to clipboard
| Challenge: | General-purpose Language Models have changed the world of Natural Language Processing, if not the world itself. |
| Approach: | This tutorial will lay the foundations and explain the basics of evaluation and compare traditional methods to newly developed methods. |
| Outcome: | The tutorial assumes little familiarity with metrics, datasets, prompts and benchmarks . it will compare traditional methods to newly developed methods . |
Mediocrity is the key for LLM as a Judge Anchor Selection (2026.acl-long)
Copied to clipboard
| Challenge: | a poor selection of an anchor can dramatically reduce correlation with human rankings . traditional reference-based metrics are often ill-suited for open-ended generation . |
| Approach: | They evaluate 22 different anchors on a Arena-Hard-v2.0 dataset and quantify the effect size of anchor selection. |
| Outcome: | The proposed model is better or worse than all other models, but it is rarely indicative of the relative ranking of the models. |
Evaluating Large Language Models on Controlled Generation Tasks (2023.emnlp-main)
Copied to clipboard
Jiao Sun, Yufei Tian, Wangchunshu Zhou, Nan Xu, Qian Hu, Rahul Gupta, John Wieting, Nanyun Peng, Xuezhe Ma
| Challenge: | Recent studies have looked into the ability of large language models in various benchmark tasks, including question generation, reading comprehension, multilingual and etc. However, few studies investigate the controllability of large languages. |
| Approach: | They propose to compare large language models with state-of-the-start finetuned smaller models to find that large language model controls are comparable to smaller models. |
| Outcome: | The proposed model can meet hard constraints and perform better than state-of-the-art models. |
LaMP: When Large Language Models Meet Personalization (2024.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for personalization in large language models are understudied . |
| Approach: | They propose a benchmark for training and evaluating language models for producing personalized outputs using a set of seven personalized tasks . they propose two retrieval augmentation approaches that retrieve personal items from each user profile for personalizing language model outputs. |
| Outcome: | The proposed approach is effective for a set of zero-shot and fine-tuned language models and highlights the impact of personalization in various natural language tasks. |