Benchmark Transparency: Measuring the Impact of Data on Evaluation (2024.naacl-long)
Copied to clipboard
| Challenge: | In this paper, we quantify the impact that data distribution has on the performance and evaluation of NLP models. |
| Approach: | They propose to use disproportional stratified sampling to measure the data distribution across 6 different dimensions to quantify model performance. |
| Outcome: | The proposed framework measures the data distribution across 6 different dimensions and shows that it is statistically significant and predicts model performance. |
Similar Papers
Privacy Evaluation Benchmarks for NLP Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Several kinds of privacy attacks are studied in depth, but they are non-systematic and lack a comprehensive understanding of the impact caused by the attacks. |
| Approach: | They propose a privacy attack and defense evaluation benchmark in the field of NLP . they propose an improved attack method and a chained framework for privacy attacks . |
| Outcome: | The proposed framework can be chained to achieve a higher-level attack objective. |
We Need to Measure Data Diversity in NLP — Better and Broader (2025.emnlp-main)
Copied to clipboard
| Challenge: | Language models exhibit remarkable natural language understanding and generation capabilities, but they have serious flaws, such as societal biases and spurious correlations. |
| Approach: | They argue that interdisciplinary perspectives are essential for developing more fine-grained and valid measures of data diversity. |
| Outcome: | The proposed measures are based on interdisciplinary perspectives and include a variety of datasets. |
NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for evaluating large language models using annotated benchmarks are in trouble . data contamination can cause wrong scientific conclusions being published . |
| Approach: | They argue that the evaluation of NLP tasks using annotated benchmarks is in trouble . they define different levels of data contamination and propose a community effort . |
| Outcome: | The proposed measures should detect when data from a benchmark was exposed to a model and flag papers with conclusions compromised by data contamination. |
Modeling Disclosive Transparency in NLP Application Descriptions (2021.emnlp-main)
Copied to clipboard
| Challenge: | Broader disclosive transparency is difficult to define and quantify, authors say . previous work has demonstrated trade-offs and negative consequences to disclosing transparency . |
| Approach: | They propose to use neural language model-based probabilistic metrics to model disclosive transparency . they demonstrate that they correlate with user and expert opinions of system transparency a valid objective proxy . |
| Outcome: | The proposed metrics correlate with user and expert opinions of system transparency, making them a valid objective proxy. |
Beyond the Tip of the Iceberg: Assessing Coherence of Text Classifiers (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Large-scale, pre-trained language models achieve human-level and superhuman accuracy on existing language understanding tasks, but statistical bias in benchmark data and probing studies has recently called into question their true capabilities. |
| Approach: | They propose to evaluate systems through a measure of prediction coherence by using two existing language understanding benchmarks with different properties to demonstrate its versatility. |
| Outcome: | The proposed evaluation framework is quick, effective, and versatile to provide insight into the coherence of machines’ predictions. |
Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets (2021.acl-long)
Copied to clipboard
| Challenge: | Several recent efforts have focused on benchmark datasets consisting of pairs of contrastive sentences, which are often accompanied by metrics that aggregate an NLP system’s behavior on these pairs into measurements of harms. |
| Approach: | They apply a measurement modeling lens to inventory pitfalls that threaten benchmarks' validity as measurement models for stereotyping. |
| Outcome: | The proposed benchmarks lack clarity and assumptions that affect how they conceptualize and operationalize stereotyping. |
Dynabench: Rethinking Benchmarking in NLP (2021.naacl-main)
Copied to clipboard
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, Adina Williams
| Challenge: | Dynabench is an open-source platform for dynamic dataset creation and model benchmarking. |
| Approach: | They propose an open-source platform for dynamic dataset creation and model benchmarking. |
| Outcome: | The proposed platform can be used to create models that fail on simple challenges and falter in real-world scenarios. |
Investigating Data Variance in Evaluations of Automatic Machine Translation Metrics (2022.findings-acl)
Copied to clipboard
| Challenge: | Current evaluation methods focus on one dataset, e.g., Newstest dataset in each year’s WMT Metrics Shared Task. |
| Approach: | They propose to use a single dataset to evaluate the performance of automatic translation metrics. |
| Outcome: | The results show that the rankings of metrics vary when the evaluation is conducted on different datasets. |
Exploring Predictive Uncertainty and Calibration in NLP: A Study on the Impact of Method & Data Scarcity (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Using low-resource languages, we assess the quality of uncertainty estimates from a wide array of approaches, but with more data. |
| Approach: | They train models on sub-sampled datasets in three different languages to assess the confidence of a neural classifier. |
| Outcome: | The proposed models train on sub-sampled datasets in three different languages and show that the quality of uncertainty estimates suffers with more data. |
What Will it Take to Fix Benchmarking in Natural Language Understanding? (2021.naacl-main)
Copied to clipboard
| Challenge: | Evaluation for many natural language understanding (NLU) tasks is broken due to unreliable and biased systems scoring so high on standard benchmarks. |
| Approach: | They argue that current benchmarks fail at four criteria for evaluation . they argue that adversarial data collection does not address the causes of failures . |
| Outcome: | The proposed frameworks fail at four criteria, and adversarial data collection does not address the causes of these failures, the authors argue . restoring a healthy evaluation ecosystem will require significant progress in the design of benchmark datasets, reliability with which they are annotated, their size, and the ways they handle social bias. |