Papers by Nageen Himayat
Soft Token Attacks Cannot Reliably Audit Unlearning in Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Recent work shows that soft token attacks can extract unlearned information from large language models. |
| Approach: | They show that soft token attacks can extract unlearned information from LLMs . |
| Outcome: | The proposed attacks can extract unlearned information from large language models . |