Papers by Justin Xu

9 papers
RadEval: A framework for radiology text evaluation (2025.emnlp-demos)

Copied to clipboard

Challenge: Evaluating automated radiology report generation systems remains a fundamental challenge in the development of safe, accurate, and clinically useful medical AI.
Approach: They propose a unified, open-source framework for evaluating radiology texts that consolidates a diverse range of metrics from classic ngram overlap (BLEU) and contextual measures (BERTScore) to clinical concept-based scores (GREEN).
Outcome: The framework consolidates a diverse range of metrics from ngram overlap (BLEU) and contextual measures (BERTScore) to clinical concept-based scores (F1CheXbert, F1RadGraph, RaTEScore, SRR-BERT, TemporalEntityF1) and advanced LLMbased evaluators (GREEN).
GREEN: Generative Radiology Report Evaluation and Error Notation (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing automated evaluation metrics fail to consider factual correctness or are limited in their interpretability.
Approach: They propose a radiology report evaluation metric that leverages natural language understanding of language models to identify and explain clinically significant errors.
Outcome: The proposed method demonstrates higher correlation with expert error counts and higher alignment with expert preferences when compared to previous methods.
Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors (2023.acl-long)

Copied to clipboard

Challenge: Abstractive summarization systems still include factual errors in generated summaries despite recent improvements in factuality detection .
Approach: They aggregate factuality error annotations from nine existing datasets and stratify them according to the underlying summarization model.
Outcome: The proposed method improves on the ChatGPT-based model and shows that it is not superior for all error types.
Morphological Segmentation for Low Resource Languages (2020.lrec-1)

Copied to clipboard

Challenge: a new corpus of annotated morphological data is described for the DARPA LORELEI Program . the data is annotating 9 low resource languages and root information for 7 of the languages .
Approach: This paper describes a new morphology resource created by Linguistic Data Consortium and the University of Pennsylvania for the DARPA LORELEI Program.
Outcome: The annotated corpus provides a gold standard for unsupervised morphological segmenters and analyzers . the language-specific annotation guidelines were language-independent, but included morphology paradigms and other specifications.
Automated Structured Radiology Report Generation (2025.acl-long)

Copied to clipboard

Challenge: Existing models struggle to produce consistent, clinically meaningful reports and standard evaluation metrics fail to capture the nuances of radiological interpretation.
Approach: They propose to reformulate free-text radiology reports into a standardized format, ensuring clarity, consistency, and structured clinical reporting.
Outcome: The proposed task reformulates free-text radiology reports into a standardized format, ensuring clarity, consistency, and structured clinical reporting.
CheXalign: Preference fine-tuning in chest X-ray interpretation models without human feedback (2025.acl-long)

Copied to clipboard

Challenge: Radiologists are a crucial role in translating medical images into actionable reports . however, the field faces staffing shortages and increasing workloads .
Approach: They propose an automated pipeline for preference feedback focusing on chest X-ray radiology report generation (RRG) method leverages publicly available datasets containing pairs of images and radiologist-written reference reports with reference-based metrics, or Judges.
Outcome: The proposed pipeline achieves state-of-the-art CheXbert scores on the MIMIC-CXR dataset while on average maintaining robust performance across six additional image perception and reasoning tasks.
LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are highly sensitive to prompts, but most automatic prompt optimization methods assume access to ground-truth references that are costly to obtain.
Approach: They propose a sample-efficient framework for label-free prompt optimization based on pairwise preference feedback from an LLM judge.
Outcome: Experiments on BIG-bench Hard and MS MARCO show that the proposed framework identifies stronger prompts than label-free baselines while offering favorable quality–cost trade-offs.
Tree-of-Quote Prompting Improves Factuality and Attribution in Multi-Hop and Medical Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) produce fluent but factually incorrect outputs, a phenomenon commonly referred to as hallucination.
Approach: They propose a Tree-of-Quote framework that decomposes complex questions into subquestions and generates quotes to support each step without retrieval.
Outcome: Experiments on StrategyQA, 2WikiMultiHopQA, MuSiQue, MoreHopQ, and MedQA show that ToQ improves factuality and attribution over baselines.
Instructional Agents: Reducing Teaching Faculty Workload through Multi-Agent Instructional Design (2026.eacl-long)

Copied to clipboard

Challenge: Existing AI-assisted educational tools focus on isolated tasks, but lack end-to-end workflows for instructional design.
Approach: They propose a multi-agent large language model framework to automate end-to-end course material generation.
Outcome: The proposed framework reduces development time and human workload while reducing human involvement.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations