Papers by Ehud Reiter

12 papers
Consultation Checklists: Standardising the Human Evaluation of Medical Note Generation (2022.emnlp-industry)

Copied to clipboard

Challenge: Existing studies have shown that note generation is difficult due to subjective nature of many aspects of output quality.
Approach: They propose a protocol that aims to increase objectivity by grounding evaluations in Consultation Checklists, which are created in a preliminary step and then used as a common point of reference during quality assessment.
Outcome: The proposed protocol shows that the evaluations produced in the study are more objective than the original human note.
Evolving Stances on Reproducibility: A Longitudinal Study of NLP and ML Researchers’ Views and Experience of Reproducibility (2025.findings-emnlp)

Copied to clipboard

Challenge: Identical experiments producing different results can be due to variation between samples of evaluation items or evaluators, but it can also be due . poor experimental practice can be mitigated by bringing multiple comparable studies together in systematic reviews that draw conclusions beyond the level of the individual studies.
Approach: They propose to assess NLP/ML practitioners' views and experience of reproducibility over the past two years.
Outcome: The results of two identical surveys show that views and experience of reproducibility have changed over the past two years.
Non-Repeatable Experiments and Non-Reproducible Results: The Reproducibility Crisis in Human Evaluation in NLP (2023.findings-acl)

Copied to clipboard

Challenge: reproducibility of human evaluations is rarely queried in NLP . authors estimate that just 5% of humanevaluations are repeatable .
Approach: They propose to make human evaluations more repeatable and more reproducible . they estimate that just 5% of human evaluation experiments are repeatable .
Outcome: The results show that human evaluations are rarely queried or formally tested in NLP . the authors estimate that just 5% of human evaluation experiments are repeatable .
Human Evaluation and Correlation with Automatic Metrics in Consultation Note Generation (2022.acl-long)

Copied to clipboard

Challenge: Recent studies suggest that note generation systems can be used to generate clinical consultation notes from the verbatim transcript of the consultation.
Approach: They propose to use machine learning to generate consultation notes from the verbatim transcript of the consultation to evaluate their effectiveness.
Outcome: The proposed model performs better than common model-based metrics like BertScore and is open-sourced.
CausalGraphBench: a Benchmark for Evaluating Language Models capabilities of Causal Graph discovery (2025.acl-srw)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have expanded their applications into domains not traditionally associated with natural language processing.
Approach: They propose a benchmark to evaluate the ability of large language models to construct Causal Graphs (CGs) they examine various methods for CG discovery and their performance across different graph sizes and complexity levels.
Outcome: The proposed benchmark comprises 35 CGs sourced from publicly available repositories and academic papers.
Improving Factual Accuracy of Neural Table-to-Text Output by Addressing Input Problems in ToTTo (2024.naacl-long)

Copied to clipboard

Challenge: Neural Table-to-Text models produce hallucinated outputs that are factually incorrect or unrelated to the input data.
Approach: They manually annotated 1,837 texts generated by multiple Neural Table-to-Text models in the politics domain of the ToTTo dataset.
Outcome: The proposed model reduces factual errors by 52% to 76% . the proposed model also struggles with tabular inputs that are structured in a non-standard way, especially when the input lacks distinct row and column values or the column headers are not correctly mapped to corresponding values.
Are Experts Needed? On Human Evaluation of Counselling Reflection Generation (2023.acl-long)

Copied to clipboard

Challenge: Language models have been used to generate reflections automatically, but human evaluation is challenging due to the cost of hiring experts.
Approach: They ask laypeople and experts to annotate synthetic reflections and human reflections from actual therapists.
Outcome: The proposed method shows that laypeople and experts are reliable annotators and have moderate-to-strong inter-group correlation.
User-Driven Research of Medical Note Generation Software (2022.naacl-main)

Copied to clipboard

Challenge: Existing studies on how NLP systems could be used in clinical practice focus on technical difficulties and usability challenges involved in implementing them.
Approach: They propose to use Speech Recognition to transcribe the audio of a medical consultation and then to train sequence-to-sequence models to summarise the transcript into a consultation note.
Outcome: The proposed system generates notes in real time during a doctor-patient consultation and is able to capture the salient points of a consultation . the proposed system is based on three rounds of user studies in a live telehealth clinic and identifies a number of clinical use cases that could prove challenging for the system.
SPHERE: An Evaluation Card for Human-AI Systems (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluation methods and standards for human-AI systems are unclear, especially for large language models.
Approach: They propose an evaluation card SPHERE which provides a template for evaluation protocols . they outline current evaluation practices and areas for improvement .
Outcome: The evaluation card provides a template for designing evaluation protocols . it outlines current evaluation practices and areas for improvement .
Scalability of Bayesian Network Structure Elicitation with Large Language Models: a Novel Methodology and Comparative Analysis (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for BN structure learning are limited by the size of the BN.
Approach: They propose a method for Bayesian Networks (BNs) structure elicitation that initializes several LLMs with different experiences and queries them to create a structure.
Outcome: The proposed method performs better than the existing method with one of the three studied LLMs, but the performance decreases with the increase in BN size.
A Systematic Review of Reproducibility Research in Natural Language Processing (2021.eacl-main)

Copied to clipboard

Challenge: Despite the recent progress in reproducibility, the field is far from reaching a consensus on how reproducibility should be defined, measured and addressed.
Approach: They propose to provide a wide-angle snapshot of current work on reproducibility in NLP.
Outcome: The proposed work will provide a wide-angle snapshot of current work on reproducibility in NLP.
Ask the experts: sourcing a high-quality nutrition counseling dataset through Human-AI collaboration (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are being used by end-users for various tasks, including sensitive ones such as health counseling, disregarding potential safety concerns.
Approach: They use ChatGPT to crowd-source dietary struggles and work with nutrition experts to generate supportive text using ChatGPS.
Outcome: The proposed model outperforms other models on dietary struggles and mental health tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations