Papers by Craig Thomson

5 papers
Standard Quality Criteria Derived from Current NLP Evaluations for Guiding Evaluation Design and Grounding Comparability and AI Compliance Assessments (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluations do not evaluate the same aspect of quality, resulting in unclear comparability and low repeatability.
Approach: They propose to use a standard set of qualitycriterion names and definitions to establish comparability of existing evaluations.
Outcome: The proposed taxonomy combines 114 quality criteria from 3 surveys of 933 evaluations in NLP and is used to establish comparability of existing evaluations and guide the design of new evaluations.
Evolving Stances on Reproducibility: A Longitudinal Study of NLP and ML Researchers’ Views and Experience of Reproducibility (2025.findings-emnlp)

Copied to clipboard

Challenge: Identical experiments producing different results can be due to variation between samples of evaluation items or evaluators, but it can also be due . poor experimental practice can be mitigated by bringing multiple comparable studies together in systematic reviews that draw conclusions beyond the level of the individual studies.
Approach: They propose to assess NLP/ML practitioners' views and experience of reproducibility over the past two years.
Outcome: The results of two identical surveys show that views and experience of reproducibility have changed over the past two years.
Non-Repeatable Experiments and Non-Reproducible Results: The Reproducibility Crisis in Human Evaluation in NLP (2023.findings-acl)

Copied to clipboard

Challenge: reproducibility of human evaluations is rarely queried in NLP . authors estimate that just 5% of humanevaluations are repeatable .
Approach: They propose to make human evaluations more repeatable and more reproducible . they estimate that just 5% of human evaluation experiments are repeatable .
Outcome: The results show that human evaluations are rarely queried or formally tested in NLP . the authors estimate that just 5% of human evaluation experiments are repeatable .
GEMv2: Multilingual NLG Benchmarking in a Single Line of Code (2022.emnlp-demos)

Copied to clipboard

Challenge: Evaluations in machine learning rarely use the latest metrics, datasets, or human evaluation in favor of remaining compatible with prior work.
Approach: They propose to use the Generation, Evaluation, and Metrics Benchmark to integrate new evaluation methods into existing evaluations.
Outcome: The proposed evaluation infrastructure bridges the gap between the advantages of leaderboards and in-depth and evolving evaluations by allowing model developers to benefit from each other's work.
Automatic Paper Analysis and Categorisation for Systematic Reviews with Combined Reasoning-Augmented SFT and DAPO RL (2026.findings-acl)

Copied to clipboard

Challenge: Automating systematic reviews is expensive and time consuming, a study finds . automatic approaches are being explored but their performance has been poor .
Approach: They propose to use reasoning-enhanced fine-tuning and DAPO reinforcement learning to automate systematic reviews.
Outcome: The proposed methods significantly improve the performance of LLMs, the authors find . they find that reasoning-enhanced fine-tuning reduces time required for annotation by 80% .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations