| Challenge: | Existing methods for comparing DNNs on unseen data are not suitable for this task. |
| Approach: | They propose to adapt a test for the Almost Stochastic Dominance relation between two distributions to the problem by comparing their performance on unseen data. |
| Outcome: | The proposed method meets all criteria while previously proposed methods fail to do so. |
Similar Papers
Towards Robust Comparisons of NLP Models: A Case Study (2025.coling-main)
Copied to clipboard
| Challenge: | Existing statistical tests to compare the test scores of different NLP models have been proposed to account for nuisance factors such as noise, randomness, or hyperparameter values. |
| Approach: | They propose a regression analysis which isolates the effect of nuisance factors from the effects of the models’ capabilities. |
| Outcome: | The proposed model is able to show that the difference between BioLinkBERT and MSR BiomedBERT is 7 times smaller than previously reported. |
Predicting generalization performance with correctness discriminators (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models estimate accuracy of models on unlabeled test data, but they hide their own uncertainty. |
| Approach: | They propose a model that establishes upper and lower bounds on the accuracy without requiring gold labels for the unseen data. |
| Outcome: | The proposed model establishes upper and lower bounds on accuracy without requiring gold labels for the unseen data. |
Proceedings of the 2nd Workshop on Deep Learning Approaches for Low-Resource NLP (DeepLo 2019) (D19-61)
Copied to clipboard
| Challenge: | EMNLP-IJCNLP 2019 Workshop on Deep Learning Approaches for Low-Resource Natural Language Processing takes place in Hong Kong, China . |
| Approach: | EMNLP-IJCNLP 2019 Workshop on Deep Learning Approaches for Low-Resource Natural Language Processing takes place in Hong Kong, China . call for papers for this second workshop met with a strong response . |
| Outcome: | the EMNLP-IJCNLP 2019 workshop on deep learning approaches for low-resource natural language processing takes place in Hong Kong, China. |
An Empirical Comparison of Instance Attribution Methods for NLP (2021.naacl-main)
Copied to clipboard
| Challenge: | Influence functions provide machinery for identifying training instances that may have led to a specific prediction, but are computationally expensive and prohibitive in many cases. |
| Approach: | They evaluate the degree to which different potential instance attribution agrees with respect to the importance of training samples. |
| Outcome: | The proposed methods exhibit desirable characteristics similar to more complex methods, but are computationally expensive. |
Confidence and Stability of Global and Pairwise Scores in NLP Evaluation (2025.acl-srw)
Copied to clipboard
| Challenge: | Modern natural language processing benchmarks are often represented as pairwise comparison leaderboards, such as LMSYS Arena. |
| Approach: | They investigate the strengths and weaknesses of global scores and pairwise comparisons to aid decision-making in selecting appropriate model evaluation strategies. |
| Outcome: | The proposed method underestimates strong models with rare errors or low confidence, while relying on global scores can be more effective. |
Even the Simplest Baseline Needs Careful Re-investigation: A Case Study on XML-CNN (2022.naacl-main)
Copied to clipboard
| Challenge: | XML-CNN has been a popular research topic in NLP due to its superior performance . however, the increasing complexity brings difficulties to ensure the true architectural progress . |
| Approach: | They propose to re-examine an influential multi-label text classification method . they propose suitable baselines for multi-level text classification tasks . |
| Outcome: | The proposed method performs better than the original model, the authors show . they show that the re-implementation reveals contradictory results to the original work . |
Neuron-level Interpretation of Deep NLP Models: A Survey (2022.tacl-1)
Copied to clipboard
| Challenge: | Existing work on deep neural networks has focused on representation analysis, but recent work focused on analyzing neurons within these models. |
| Approach: | They propose to analyze neural networks to uncover linguistic concepts captured by the network . they propose to use a granular approach to analyze neurons within these models . |
| Outcome: | The proposed method combines methods to discover and understand neurons in a network with evaluation methods. |
A Survey of Methods for Addressing Class Imbalance in Deep-Learning Based Natural Language Processing (2023.eacl-main)
Copied to clipboard
| Challenge: | Developing methods to improve model performance in imbalanced data settings has been an active area for decades . |
| Approach: | They propose to use sampling, data augmentation, choice of loss function, staged learning, or model design to address class imbalance in NLP. |
| Outcome: | The proposed approaches are evaluated on a variety of NLP tasks or in the computer vision community. |
Neural Reranking for Dependency Parsing: An Evaluation (2020.acl-main)
Copied to clipboard
| Challenge: | Recent work shows that neural rerankers can improve dependency parsing results over the top k trees produced by a base parser. |
| Approach: | They propose to use a discriminative reranker to improve dependency parsing results . they propose to incorporate global information into the model to improve parse accuracies . |
| Outcome: | The proposed model outperforms existing models on English and German and Czech, and is the only one to improve on German and Chinese data. |
Please, Don’t Forget the Difference and the Confidence Interval when Seeking for the State-of-the-Art Status (2022.lrec-1)
Copied to clipboard
| Challenge: | comparing NLP systems by performance has become an essential question . comparing systems by performing performance criterion is criticized for allowing chance to determine superiority . |
| Approach: | They propose to use bootstrap confidence intervals instead of state-of-the-art status and statistical significance testing to compare NLP system performance. |
| Outcome: | The bootstrap confidence intervals are used to compare NLP system performance . the bootstrap test is more accurate than state-of-the-art status and statistical significance testing . |