Challenge: Existing methods for comparing DNNs on unseen data are not suitable for this task.
Approach: They propose to adapt a test for the Almost Stochastic Dominance relation between two distributions to the problem by comparing their performance on unseen data.
Outcome: The proposed method meets all criteria while previously proposed methods fail to do so.

Similar Papers

Towards Robust Comparisons of NLP Models: A Case Study (2025.coling-main)

Copied to clipboard

Challenge: Existing statistical tests to compare the test scores of different NLP models have been proposed to account for nuisance factors such as noise, randomness, or hyperparameter values.
Approach: They propose a regression analysis which isolates the effect of nuisance factors from the effects of the models’ capabilities.
Outcome: The proposed model is able to show that the difference between BioLinkBERT and MSR BiomedBERT is 7 times smaller than previously reported.
Predicting generalization performance with correctness discriminators (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing models estimate accuracy of models on unlabeled test data, but they hide their own uncertainty.
Approach: They propose a model that establishes upper and lower bounds on the accuracy without requiring gold labels for the unseen data.
Outcome: The proposed model establishes upper and lower bounds on accuracy without requiring gold labels for the unseen data.
Proceedings of the 2nd Workshop on Deep Learning Approaches for Low-Resource NLP (DeepLo 2019) (D19-61)

Copied to clipboard

Challenge: EMNLP-IJCNLP 2019 Workshop on Deep Learning Approaches for Low-Resource Natural Language Processing takes place in Hong Kong, China .
Approach: EMNLP-IJCNLP 2019 Workshop on Deep Learning Approaches for Low-Resource Natural Language Processing takes place in Hong Kong, China . call for papers for this second workshop met with a strong response .
Outcome: the EMNLP-IJCNLP 2019 workshop on deep learning approaches for low-resource natural language processing takes place in Hong Kong, China.
An Empirical Comparison of Instance Attribution Methods for NLP (2021.naacl-main)

Copied to clipboard

Challenge: Influence functions provide machinery for identifying training instances that may have led to a specific prediction, but are computationally expensive and prohibitive in many cases.
Approach: They evaluate the degree to which different potential instance attribution agrees with respect to the importance of training samples.
Outcome: The proposed methods exhibit desirable characteristics similar to more complex methods, but are computationally expensive.
Confidence and Stability of Global and Pairwise Scores in NLP Evaluation (2025.acl-srw)

Copied to clipboard

Challenge: Modern natural language processing benchmarks are often represented as pairwise comparison leaderboards, such as LMSYS Arena.
Approach: They investigate the strengths and weaknesses of global scores and pairwise comparisons to aid decision-making in selecting appropriate model evaluation strategies.
Outcome: The proposed method underestimates strong models with rare errors or low confidence, while relying on global scores can be more effective.
Even the Simplest Baseline Needs Careful Re-investigation: A Case Study on XML-CNN (2022.naacl-main)

Copied to clipboard

Challenge: XML-CNN has been a popular research topic in NLP due to its superior performance . however, the increasing complexity brings difficulties to ensure the true architectural progress .
Approach: They propose to re-examine an influential multi-label text classification method . they propose suitable baselines for multi-level text classification tasks .
Outcome: The proposed method performs better than the original model, the authors show . they show that the re-implementation reveals contradictory results to the original work .
Neuron-level Interpretation of Deep NLP Models: A Survey (2022.tacl-1)

Copied to clipboard

Challenge: Existing work on deep neural networks has focused on representation analysis, but recent work focused on analyzing neurons within these models.
Approach: They propose to analyze neural networks to uncover linguistic concepts captured by the network . they propose to use a granular approach to analyze neurons within these models .
Outcome: The proposed method combines methods to discover and understand neurons in a network with evaluation methods.
A Survey of Methods for Addressing Class Imbalance in Deep-Learning Based Natural Language Processing (2023.eacl-main)

Copied to clipboard

Challenge: Developing methods to improve model performance in imbalanced data settings has been an active area for decades .
Approach: They propose to use sampling, data augmentation, choice of loss function, staged learning, or model design to address class imbalance in NLP.
Outcome: The proposed approaches are evaluated on a variety of NLP tasks or in the computer vision community.
Neural Reranking for Dependency Parsing: An Evaluation (2020.acl-main)

Copied to clipboard

Challenge: Recent work shows that neural rerankers can improve dependency parsing results over the top k trees produced by a base parser.
Approach: They propose to use a discriminative reranker to improve dependency parsing results . they propose to incorporate global information into the model to improve parse accuracies .
Outcome: The proposed model outperforms existing models on English and German and Czech, and is the only one to improve on German and Chinese data.
Please, Don’t Forget the Difference and the Confidence Interval when Seeking for the State-of-the-Art Status (2022.lrec-1)

Copied to clipboard

Challenge: comparing NLP systems by performance has become an essential question . comparing systems by performing performance criterion is criticized for allowing chance to determine superiority .
Approach: They propose to use bootstrap confidence intervals instead of state-of-the-art status and statistical significance testing to compare NLP system performance.
Outcome: The bootstrap confidence intervals are used to compare NLP system performance . the bootstrap test is more accurate than state-of-the-art status and statistical significance testing .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations