Papers by Jon Burnsky
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization (2024.naacl-long)
Copied to clipboard
Liyan Tang, Igor Shalyminov, Amy Wong, Jon Burnsky, Jake Vincent, Yu’an Yang, Siffi Singh, Song Feng, Hwanjun Song, Hang Su, Lijia Sun, Yi Zhang, Saab Mansour, Kathleen McKeown
| Challenge: | Existing LLMs hallucinate significant amounts of factual errors in the dialogue domain, regardless of the model’s size. |
| Approach: | They propose to evaluate topic-focused dialogue summarization by using large language models (LLMs) they use human annotations to evaluate factual consistency and explain factually inconsistent sentences. |
| Outcome: | The proposed evaluation benchmark on topic-focused dialogue summarization shows that existing LLMs hallucinate significant amounts of factual errors regardless of the model’s size. |
TN-Eval: Rubric and Evaluation Protocols for Measuring the Quality of Behavioral Therapy Notes (2025.acl-industry)
Copied to clipboard
| Challenge: | Behavioral therapy notes are important for legal compliance and patient care, but quality standards for them remain underdeveloped. |
| Approach: | They propose a rubric for evaluating therapy notes across key dimensions: completeness, conciseness, faithfulness. |
| Outcome: | The proposed evaluation framework improves on therapist-written notes and LLM-generated notes. |