Challenge: a lack of reproducibility and generalisability is a major threat to scientific development in Natural Language Processing.
Approach: They propose to use a model zoo to document and release language models and published code . they recommend that future replication experiments should consider a variety of datasets .
Outcome: The proposed methods are compared on six English datasets and are based on the results.

Similar Papers

A Systematic Review of Reproducibility Research in Natural Language Processing (2021.eacl-main)

Copied to clipboard

Challenge: Despite the recent progress in reproducibility, the field is far from reaching a consensus on how reproducibility should be defined, measured and addressed.
Approach: They propose to provide a wide-angle snapshot of current work on reproducibility in NLP.
Outcome: The proposed work will provide a wide-angle snapshot of current work on reproducibility in NLP.
Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and Future (2023.emnlp-main)

Copied to clipboard

Challenge: Existing literature on the generalization of machine learning models to out-of-distribution data is lacking.
Approach: They propose to present the first comprehensive review of recent progress, methods, and evaluations on the generalization challenge from an OOD perspective in natural language understanding.
Outcome: The proposed survey provides the first comprehensive review of recent progress, methods, and evaluations on the generalization challenge from an OOD perspective in natural language understanding.
NLP Reproducibility For All: Understanding Experiences of Beginners (2023.acl-long)

Copied to clipboard

Challenge: a study with 93 students in an introductory NLP class shows that beginners' programming skill and comprehension of research papers have a limited impact on their effort spent completing the exercise.
Approach: a study conducted with 93 students in an introductory NLP course questioned them on their programming background and programming background.
Outcome: The results show that beginners' programming skill and comprehension of research papers have a limited impact on their effort spent on the exercise.
Principles from Clinical Research for NLP Model Generalization (2024.naacl-long)

Copied to clipboard

Challenge: In clinical research, generalizability depends on (a) internal validity of experiments and (b) external validity or transportability of the results to the wider population.
Approach: They propose to ensure internal validity when building machine learning models in NLP by incorporating learning spurious correlations into their models.
Outcome: The proposed model can perform well on data unseen during training, but drawn from the same distribution or population.
On the Gap between Adoption and Understanding in NLP (2021.findings-acl)

Copied to clipboard

Challenge: a recent paper argues that current publications foster a gap between adoption and understanding of models . it also makes it easier to meet publication demands with method papers, argues the paper .
Approach: They argue that current NLP publication models foster a gap between adoption and understanding of models . they argue that everlarger models make it harder to explain how our methods work .
Outcome: The authors argue that current publications foster a gap between adoption and understanding of models . they argue that the rise of everlarger models makes it harder to explain how our methods work .
Non-Repeatable Experiments and Non-Reproducible Results: The Reproducibility Crisis in Human Evaluation in NLP (2023.findings-acl)

Copied to clipboard

Challenge: reproducibility of human evaluations is rarely queried in NLP . authors estimate that just 5% of humanevaluations are repeatable .
Approach: They propose to make human evaluations more repeatable and more reproducible . they estimate that just 5% of human evaluation experiments are repeatable .
Outcome: The results show that human evaluations are rarely queried or formally tested in NLP . the authors estimate that just 5% of human evaluation experiments are repeatable .
Reproduction and Replication: A Case Study with Automatic Essay Scoring (2020.lrec-1)

Copied to clipboard

Challenge: reproducibility of experiments has gained more attention in the NLP community . recent negative reproduction results indicate that published results are not verifiable .
Approach: They propose to reproduce an earlier study of automatic essay scoring for determining the proficiency of second language learners in a multilingual setting.
Outcome: The proposed reproduction of an AES system for determining the proficiency of second language learners in a multilingual setting is compared with the original.
Evolving Stances on Reproducibility: A Longitudinal Study of NLP and ML Researchers’ Views and Experience of Reproducibility (2025.findings-emnlp)

Copied to clipboard

Challenge: Identical experiments producing different results can be due to variation between samples of evaluation items or evaluators, but it can also be due . poor experimental practice can be mitigated by bringing multiple comparable studies together in systematic reviews that draw conclusions beyond the level of the individual studies.
Approach: They propose to assess NLP/ML practitioners' views and experience of reproducibility over the past two years.
Outcome: The results of two identical surveys show that views and experience of reproducibility have changed over the past two years.
A Major Obstacle for NLP Research: Let’s Talk about Time Allocation! (2022.emnlp-main)

Copied to clipboard

Challenge: Subpar time allocation has been a major obstacle for natural language processing research in recent years, argues a new position paper .
Approach: They propose to identify the biggest traps the NLP community falls into and suggest solutions to solve them.
Outcome: The authors outline multiple concrete problems together with their negative consequences and suggest remedies to improve the status quo.
From Annotation to Adaptation: Metrics, Synthetic Data, and Aspect Extraction for Aspect-Based Sentiment Analysis with Large Language Models (2025.naacl-srw)

Copied to clipboard

Challenge: Using a synthetic sports feedback dataset, we evaluate open-weight LLMs’ ability to extract aspect-polarity pairs.
Approach: They propose a metric to facilitate the evaluation of aspect extraction with generative models.
Outcome: The proposed metric improves the performance of open-weight LLMs in the Aspect-Based Sentiment Analysis task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations