Papers by Liyan Tang
Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding (2025.emnlp-main)
Copied to clipboard
| Challenge: | In many domains, determining faithfulness of a claim to a source document is a binary judgment . but, whether a document is factual or whether it is entailed given some input is highly subjective. |
| Approach: | They propose a task to manage the subjectivity involved with factuality judgments of ambiguous claims. |
| Outcome: | The proposed method improves the annotator agreement on faithfulness of a claim by 21%. |
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization (2024.naacl-long)
Copied to clipboard
Liyan Tang, Igor Shalyminov, Amy Wong, Jon Burnsky, Jake Vincent, Yu’an Yang, Siffi Singh, Song Feng, Hwanjun Song, Hang Su, Lijia Sun, Yi Zhang, Saab Mansour, Kathleen McKeown
| Challenge: | Existing LLMs hallucinate significant amounts of factual errors in the dialogue domain, regardless of the model’s size. |
| Approach: | They propose to evaluate topic-focused dialogue summarization by using large language models (LLMs) they use human annotations to evaluate factual consistency and explain factually inconsistent sentences. |
| Outcome: | The proposed evaluation benchmark on topic-focused dialogue summarization shows that existing LLMs hallucinate significant amounts of factual errors regardless of the model’s size. |
Understanding Factual Errors in Summarization: Errors, Summarizers, Datasets, Error Detectors (2023.acl-long)
Copied to clipboard
Liyan Tang, Tanya Goyal, Alex Fabbri, Philippe Laban, Jiacheng Xu, Semih Yavuz, Wojciech Kryscinski, Justin Rousseau, Greg Durrett
| Challenge: | Abstractive summarization systems still include factual errors in generated summaries despite recent improvements in factuality detection . |
| Approach: | They aggregate factuality error annotations from nine existing datasets and stratify them according to the underlying summarization model. |
| Outcome: | The proposed method improves on the ChatGPT-based model and shows that it is not superior for all error types. |
MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents (2024.emnlp-main)
Copied to clipboard
| Challenge: | Current methods for fact-checking are based on verifying each piece of a model against potential evidence using an LLM. |
| Approach: | They propose a method that builds small fact-checking models that have GPT-4-level performance but 400x lower cost. |
| Outcome: | The proposed model outperforms other models and reaches GPT-4 accuracy. |
Less Likely Brainstorming: Using Language Models to Generate Alternative Hypotheses (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to reduce cognitive errors in MRI interpretations do not work for generating less likely outputs. |
| Approach: | They propose a task that asks a model to generate outputs that humans think are relevant but less likely to happen. |
| Outcome: | The proposed method compares with several state-of-the-art controlled text generation models via automatic and human evaluations and shows that it reduces cognitive errors in interpreting MRI findings. |