Challenge: Existing studies show that standard splits produce low reproducible and unreliable conclusions . reproducibility of empirical experimental conclusions is a problem in NLP domain .
Approach: They propose to transform the reproducibility of a model comparison into a probabilistic function . they propose to use a regularized corpus splitting strategy to estimate the model's performance .
Outcome: The proposed estimator achieves a high SNR and significantly increases reproducibility.

Similar Papers

A Systematic Review of Reproducibility Research in Natural Language Processing (2021.eacl-main)

Copied to clipboard

Challenge: Despite the recent progress in reproducibility, the field is far from reaching a consensus on how reproducibility should be defined, measured and addressed.
Approach: They propose to provide a wide-angle snapshot of current work on reproducibility in NLP.
Outcome: The proposed work will provide a wide-angle snapshot of current work on reproducibility in NLP.
We Need to Talk about Standard Splits (P19-1)

Copied to clipboard

Challenge: Existing methods to evaluate systems with a held-out test set are insufficient for system comparison.
Approach: They propose to use multiple random splits to compare performance of systems . they replicate results on standard split but fail to reproduce some rankings .
Outcome: The proposed method is based on multiple random splits to replicate results with a set of part-of-speech taggers.
Evolving Stances on Reproducibility: A Longitudinal Study of NLP and ML Researchers’ Views and Experience of Reproducibility (2025.findings-emnlp)

Copied to clipboard

Challenge: Identical experiments producing different results can be due to variation between samples of evaluation items or evaluators, but it can also be due . poor experimental practice can be mitigated by bringing multiple comparable studies together in systematic reviews that draw conclusions beyond the level of the individual studies.
Approach: They propose to assess NLP/ML practitioners' views and experience of reproducibility over the past two years.
Outcome: The results of two identical surveys show that views and experience of reproducibility have changed over the past two years.
Split and Rephrase: Better Evaluation and Stronger Baselines (P18-2)

Copied to clipboard

Challenge: a dataset mapping a complex sentence to a sequence of sentences conveying the same meaning is challenging in NLP.
Approach: They propose a neural split and a copy-mechanism to break a complex sentence into several shorter sentences that convey the same meaning.
Outcome: The proposed model outperforms the baseline model by 8.68 BLEU and further improves on the task.
Non-Repeatable Experiments and Non-Reproducible Results: The Reproducibility Crisis in Human Evaluation in NLP (2023.findings-acl)

Copied to clipboard

Challenge: reproducibility of human evaluations is rarely queried in NLP . authors estimate that just 5% of humanevaluations are repeatable .
Approach: They propose to make human evaluations more repeatable and more reproducible . they estimate that just 5% of human evaluation experiments are repeatable .
Outcome: The results show that human evaluations are rarely queried or formally tested in NLP . the authors estimate that just 5% of human evaluation experiments are repeatable .
Reproducibility in NLP: What Have We Learned from the Checklist? (2023.findings-acl)

Copied to clipboard

Challenge: Scientific progress in NLP rests on the reproducibility of researchers’ claims.
Approach: They examine 10,405 anonymous responses to the NLP Reproducibility Checklist . they find evidence of an increase in reporting of information after the Checklist's introduction .
Outcome: The authors find that 44% of submissions that gather new data are 5% less likely to be accepted than those that did not.
NLP Reproducibility For All: Understanding Experiences of Beginners (2023.acl-long)

Copied to clipboard

Challenge: a study with 93 students in an introductory NLP class shows that beginners' programming skill and comprehension of research papers have a limited impact on their effort spent completing the exercise.
Approach: a study conducted with 93 students in an introductory NLP course questioned them on their programming background and programming background.
Outcome: The results show that beginners' programming skill and comprehension of research papers have a limited impact on their effort spent on the exercise.
Is the Best Better? Bayesian Statistical Model Comparison for Natural Language Processing (2020.emnlp-main)

Copied to clipboard

Challenge: a recent study raises concerns about the use of standard splits to compare models . we compare the performance of six English part-of-speech taggers to those of other models based on standard split analysis .
Approach: They propose a Bayesian statistical model comparison technique using k-fold cross-validation . they rank six English part-of-speech taggers across two data sets and three evaluation metrics .
Outcome: The proposed method ranks English part-of-speech taggers on two data sets and three evaluation metrics.
Reproducing and Regularizing the SCRN Model (C18-1)

Copied to clipboard

Challenge: Recurrent neural networks (RNNs) have demonstrated tremendous success in sequence modeling . naive dropout, variational dropout and weight tying are common techniques used to regularize the SCRN model .
Approach: They propose a Structurally Constrained Recurrent Network (SCRN) model and regularize it using existing techniques.
Outcome: The proposed model outperforms the LSTM model on non-English data while being much simpler.
Towards Reproducible Machine Learning Research in Natural Language Processing (2022.acl-tutorials)

Copied to clipboard

Challenge: a tutorial on reproducibility in ML addresses the problem of research results that are not reproducible.
Approach: They propose a tutorial to ensure reproducible research in ML with an emphasis on computational linguistics and NLP.
Outcome: The proposed tutorial focuses on computational linguistics and NLP . it provides a framework for using reproducibility as a teaching tool in university-level computer science programs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations