Papers by Yueqi Xie
MLLM-Protector: Ensuring MLLM’s Safety without Hurting Performance (2024.emnlp-main)
Copied to clipboard
Renjie Pi, Tianyang Han, Jianshu Zhang, Yueqi Xie, Rui Pan, Qing Lian, Hanze Dong, Jipeng Zhang, Tong Zhang
| Challenge: | MLLMs are deployed on limited image-text pairs, which makes them more vulnerable to catastrophic forgetting of their original abilities during safety fine-tuning. |
| Approach: | They propose a plug-and-play strategy that detects harmful visual inputs and transforms harmful ones into harmless ones. |
| Outcome: | The proposed approach mitigates the risks posed by malicious visual inputs without compromising the original performance of MLLMs. |
Defending against Indirect Prompt Injection by Instruction Detection (2025.findings-emnlp)
Copied to clipboard
Tongyu Wen, Chenglong Wang, Xiyuan Yang, Haoyu Tang, Yueqi Xie, Lingjuan Lyu, Zhicheng Dou, Fangzhao Wu
| Challenge: | Indirect Prompt Injection attacks can be exploited by LLMs that are embedded with external data. |
| Approach: | They propose a detection-based approach that leverages the behavioral states of LLMs to identify potential IPI attacks. |
| Outcome: | The proposed approach reduces the success rate of attacks to 0.03% on the BIPIA benchmark. |
Red-Teaming NSFW Image Classifiers as Text-to-Image Safeguards (2026.findings-acl)
Copied to clipboard
Tinghao Xie, Yueqi Xie, Alireza Zareian, Shuming Hu, Felix Juefei-Xu, Xiaowen Lin, Ankit Jain, Prateek Mittal, Li Chen
| Challenge: | Not Safe for Work (NSFW) image classifiers play a critical role in safeguarding text-to-image systems. |
| Approach: | They propose an automated red-teaming framework that leverages a set of generative AI tools to uncover NSFW image failures. |
| Outcome: | The proposed framework uncovers and interprets failure modes and enables it to be applied to real-world T2I and T2V systems. |
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods for detecting jailbreak prompts are primarily online moderation APIs or finetuned LLMs. |
| Approach: | They propose a method which scrutinizes the gradients of safety-critical parameters in large LLMs to detect jailbreak prompts. |
| Outcome: | The proposed method outperforms Llama Guard in detecting jailbreak prompts despite extensive finetuning with a large dataset. |
Measuring Human Contribution in AI-Assisted Content Generation (2026.acl-long)
Copied to clipboard
Yueqi Xie, Tao Qi, Jingwei Yi, Xiyuan Yang, Ryan Whalen, Junming Huang, Qian Ding, Yu Xie, Xing Xie, Fangzhao Wu
| Challenge: | generative AI has created a new way to generate content with humans . varying degrees of human contribution in content generation poses significant challenges for the delineation of originality . |
| Approach: | They propose a framework to measure human contribution in AI-assisted content generation by calculating mutual information between human input and AI-aided output relative to self-information of AI-assist output. |
| Outcome: | The proposed measure discriminates between varying degrees of human contribution across multiple creative domains and is validated in real-world applications. |