Papers by Wuyang Zhang
Your LLM Agent Can Leak Your Data: Data Exfiltration via Backdoored Tool Use (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing studies have not examined how backdoored agents can influence tool-use sequences to perform harmful actions. |
| Approach: | They propose a backdoor attack framework that embeds semantic triggers into fine-tuned LLM agents. |
| Outcome: | The proposed framework embeds semantic triggers into fine-tuned LLM agents . when triggered, the backdoored agent invokes memory-access tool calls to retrieve stored user context and exfiltrates it via disguised retrieval tool calls. |