Papers by Grace Byun
CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing language models fail to detect high-risk situations such as suicide ideation and child abuse . |
| Approach: | They propose a benchmark for multi-faceted mental health crisis detection that incorporates temporal labels. |
| Outcome: | The proposed benchmark significantly outperforms single-model annotations on social media posts and development examples. |
Measuring Sycophancy of Language Models in Multi-turn Dialogues (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Prior research on sycophancy has focused on single-turn factual correctness, overlooking the dynamics of real-world interactions. |
| Approach: | They propose a new evaluation suite that assesses sycophantic behavior in multi-turn, free-form conversational settings. |
| Outcome: | The proposed evaluation suite measures how quickly a model conforms to the user and how frequently it shifts its stance under sustained user pressure. |
D-GEN: Automatic Distractor Generation and Evaluation for Reliable Assessment of Generative Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for generating generative models with open-ended generation rely on predefined distractors and are costly and time-consuming. |
| Approach: | They propose a ranking alignment and entropy analysis to evaluate distractors' quality. |
| Outcome: | The proposed model preserves ranking consistency and matches the entropy distribution of ground-truth distractors. |