Papers by Yuanfan Li
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training (2025.acl-long)
Copied to clipboard
| Challenge: | Existing MGT detectors are vulnerable to simple perturbations and adversarial attacks. |
| Approach: | They propose an adversarial framework for training a robust machine-generated text detector called GREedy Adversary PromoTed DefendER. |
| Outcome: | The proposed framework reduces the Attack Success Rate (ASR) by 0.67% compared with SOTA defense methods. |