Papers by Megan Ung
Improving Model Evaluation using SMART Filtering of Benchmark Datasets (2025.naacl-long)
Copied to clipboard
| Challenge: | Creating high quality human-annotated datasets is difficult due to dataset saturation. |
| Approach: | They propose a method to filter a subset of test examples from existing benchmarks by removing less informative and lower quality examples. |
| Outcome: | The proposed method reduces dataset size by 48% while increasing Pearson correlation with rankings from ChatBot Arena. |
Training Models to Generate, Recognize, and Reframe Unhelpful Thoughts (2023.acl-long)
Copied to clipboard
| Challenge: | Existing models for cognitive behavioral therapy lack specific and diverse practice material. |
| Approach: | They propose to use a dataset to generate unhelpful thought patterns . they propose to train and evaluate existing models to generate an abundance of practice material . |
| Outcome: | The proposed model can generate unlimited quantity of practice material and generate suitable reframing proposals with no or minimal additional model training required. |
SaFeRDialogues: Taking Feedback Gracefully after Conversational Safety Failures (2022.acl-long)
Copied to clipboard
| Challenge: | Existing open-domain conversational models can easily be made to talk in inadequate ways. |
| Approach: | They propose a task and dataset of graceful responses to safety feedback . they collect 8k dialogues demonstrating safety failures, feedback signaling them, and a response acknowledging feedback. |
| Outcome: | The proposed model improves on a dataset of 8k dialogues demonstrating safety failures, feedback signaling them, and a response acknowledging the feedback. |
Learning New Skills after Deployment: Improving open-domain internet-driven dialogue with human feedback (2023.acl-long)
Copied to clipboard
| Challenge: | Frozen models trained to mimic static datasets can never improve their performance. |
| Approach: | They propose to use binary quality measurements and free-form text feedback to improve conversational skills in a conversational learning framework. |
| Outcome: | The proposed model improves on the DIRECTOR model, which is based on binary quality measurements and free-form text feedback, and shows that iterative retraining and redeployment can improve the model. |
Arbiters of Ambivalence: Challenges of using LLMs in No-Consensus tasks (2025.findings-acl)
Copied to clipboard
| Challenge: | LLMs are increasingly being used to replace humans in "aligning" LLM training . studies question this trend, but have found they can be more effective in ambivalent scenarios where humans disagree . |
| Approach: | They develop a “no-consensus” benchmark by curating examples that encompass a variety of a priori ambivalent scenarios. |
| Outcome: | The proposed benchmarks show that LLMs can provide nuanced assessments when generating open-ended answers, but tend to take a stance on no-consensus topics when employed as judges or debaters. |
ROBBIE: Robust Bias Evaluation of Large Generative Language Models (2023.emnlp-main)
Copied to clipboard
David Esiobu, Xiaoqing Tan, Saghar Hosseini, Megan Ung, Yuchen Zhang, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, Eric Smith
| Challenge: | generative large language models (LLMs) are becoming more performant and prevalent . we need tools to measure and improve their fairness, authors say . |
| Approach: | They propose to compare 6 different prompt-based bias and toxicity metrics across 12 demographic axes and 5 families of generative large language models. |
| Outcome: | The proposed model can be tested on more datasets to better characterize and mitigate biases . the study compared 6 prompt-based bias and toxicity metrics across 12 demographic axes and 5 families of generative large language models. |