Papers by Shenji Wan
Reveal and Release: Iterative LLM Unlearning with Self-generated Data (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to unlearning large language models assume full access to the forget dataset, overlooking two key challenges: (1) Forget data is often privacy-sensitive, rare, or legally regulated, making it expensive or impractical to obtain (2) The distribution of available forget data may not align with how that information is represented within the model. |
| Approach: | They propose a “Reveal-and-Release” method to unlearn with self-generated data, prompting the model to reveal what it knows using optimized instructions. |
| Outcome: | The proposed method removes the influence of undesirable data from the model. |
Inter-Passage Verification for Multi-evidence Multi-answer QA (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing multi-answer question answering systems struggle to retrieve and synthesize a large number of evidence passages. |
| Approach: | They propose a multi-answer question answering framework that generates a large set of passages and then processes each passage individually to generate an initial high-recall but noisy answer set. |
| Outcome: | The proposed framework outperforms baselines on the QAMPARI and RoMQA datasets, achieving an average F1 score improvement of 11.17%. |