The Hitchhiker’s Guide to Testing Statistical Significance in Natural Language Processing (P18-1)
Copied to clipboard
| Challenge: | Statistical significance testing is a standard statistical tool designed to ensure that experimental results are not coincidental. |
| Approach: | They propose a protocol for statistical significance test selection in NLP setups . they propose he proposes a survey of the most relevant tests to help guide the protocol . |
| Outcome: | The proposed protocol includes a survey of the most relevant tests. |
Similar Papers
Not All Claims are Created Equal: Choosing the Right Statistical Approach to Assess Hypotheses (2020.acl-main)
Copied to clipboard
| Challenge: | Empirical research in natural language processing has adopted a narrow set of principles for assessing hypotheses . alternative approaches to assess hypothese rely on p-value computation, which suffers from several known issues. |
| Approach: | They propose to compare different methods for assessing hypotheses . they argue that practitioners should first decide their target hypothesis before choosing a method . |
| Outcome: | The proposed method differs from other methods, but is not widely used in NLP . the proposed method is based on a p-value computation, but has a small gap in accuracy . |
Predicting Performance for Natural Language Processing Tasks (2020.acl-main)
Copied to clipboard
| Challenge: | Natural language processing (NLP) is a vast field, with a wide variety of tasks, languages, and domains. |
| Approach: | They build regression models to predict evaluation score of an NLP experiment . they find that their models can produce meaningful predictions over unseen languages . |
| Outcome: | The proposed model outperforms baseline models and human experts on 9 different tasks. |
NLPStatTest: A Toolkit for Comparing NLP System Performance (2020.aacl-demo)
Copied to clipboard
| Challenge: | Statistical significance testing is used to compare NLP system performance, but p-values alone are insufficient because statistical significance differs from practical significance. |
| Approach: | They propose a three-stage procedure for comparing NLP system performance and a toolkit that automates the process. |
| Outcome: | The proposed procedure is based on a three-stage procedure and compares it with existing statistical testing toolkits. |
Hard and Soft Evaluation of NLP models with BOOtSTrap SAmpling - BooStSa (2022.acl-demo)
Copied to clipboard
| Challenge: | Developing better methods for a task is a common feature of the computational linguistics literature. |
| Approach: | They propose to use bootstrap to compute significance levels with the BOOtSTrap SAmpling procedure to evaluate models that predict hard labels and soft labels as well. |
| Outcome: | The proposed method can be used to evaluate models that predict hard labels and soft labels on benchmark data sets. |
Faithful Model Evaluation for Model-Based Metrics (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies do not consider variance change due to metric model errors, which can lead to wrong conclusions. |
| Approach: | They establish the mathematical foundation of significance testing for model-based metrics . they show that metric errors can change the conclusions in certain experiments . |
| Outcome: | The proposed method can be used to derive accurate conclusions using model evaluations. |
Towards Robust Comparisons of NLP Models: A Case Study (2025.coling-main)
Copied to clipboard
| Challenge: | Existing statistical tests to compare the test scores of different NLP models have been proposed to account for nuisance factors such as noise, randomness, or hyperparameter values. |
| Approach: | They propose a regression analysis which isolates the effect of nuisance factors from the effects of the models’ capabilities. |
| Outcome: | The proposed model is able to show that the difference between BioLinkBERT and MSR BiomedBERT is 7 times smaller than previously reported. |
Causal Inference in Natural Language Processing: Estimation, Prediction, Interpretation and Beyond (2022.tacl-1)
Copied to clipboard
Amir Feder, Katherine A. Keith, Emaad Manzoor, Reid Pryzant, Dhanya Sridhar, Zach Wood-Doughty, Jacob Eisenstein, Justin Grimmer, Roi Reichart, Margaret E. Roberts, Brandon M. Stewart, Victor Veitch, Diyi Yang
| Challenge: | causality has not had the same importance in natural language processing, says aaron e. smith . he says research on causality in NLP remains scattered across domains without unified definitions . |
| Approach: | They propose to consolidate research on causality in NLP across academic areas . they explore potential uses of causal inference to improve robustness, fairness, interpretability . |
| Outcome: | The proposed method is a unified overview of causal inference for the NLP community. |
A Tutorial on Evaluation Metrics used in Natural Language Generation (2021.naacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field. |
| Approach: | This tutorial presents the evolution of automatic evaluation metrics to their current state . it aims to assess the extent of scientific progress made and identify areas/components that need improvement . |
| Outcome: | This tutorial presents the evolution of automatic evaluation metrics to their current state along with emerging trends in this field. |
Please, Don’t Forget the Difference and the Confidence Interval when Seeking for the State-of-the-Art Status (2022.lrec-1)
Copied to clipboard
| Challenge: | comparing NLP systems by performance has become an essential question . comparing systems by performing performance criterion is criticized for allowing chance to determine superiority . |
| Approach: | They propose to use bootstrap confidence intervals instead of state-of-the-art status and statistical significance testing to compare NLP system performance. |
| Outcome: | The bootstrap confidence intervals are used to compare NLP system performance . the bootstrap test is more accurate than state-of-the-art status and statistical significance testing . |
Reliability Testing for Natural Language Processing Systems (2021.acl-long)
Copied to clipboard
| Challenge: | a lack of rigorous testing and ML implicit assumption of identical training and testing distributions may result in systems that discriminate against minorities. |
| Approach: | They argue that reliability testing is needed to address the issue of demographics . they argue that adversarial attacks can be reframed for this goal . |
| Outcome: | The proposed framework will enable rigorous and targeted testing and aid in the enactment and enforcement of industry standards. |