Papers by Mazda Moayeri
DyePack: Provably Flagging Test Set Contamination in LLMs Using Backdoors (2025.emnlp-main)
Copied to clipboard
| Challenge: | Open benchmarks are essential for evaluating large language models, but their accessibility makes them likely targets of test set contamination. |
| Approach: | They propose a framework that leverages backdoor attacks to flag models that used benchmark test sets during training. |
| Outcome: | The proposed framework detects models that trained on benchmark test sets without loss of logits or internal details . it can prevent false accusations while providing strong evidence for every detected case of contamination. |