Papers by Ariel Shaulov
Safeguarding Language Models via Self-Destruct Trapdoor (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing mechanisms to restrict behavior of language models (LMs) are vulnerable to misuse and misalignment. |
| Approach: | They propose a mechanism to restrict specific behaviors in language models by exploiting hardware properties. |
| Outcome: | The proposed mechanism can be applied to trigger overflows for specific behaviors or target hardware malfunctions. |