Papers by Juan Zhai
An Optimizable Suffix Is Worth A Thousand Templates: Efficient Black-box Jailbreaking without Affirmative Phrases via LLM as Optimizer (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing jailbreaking methods generate harmful and unethical content when subjected to jailbreaking attacks. |
| Approach: | They propose a black-box jailbreaking method with optimizable suffixes that translate jailbreaking objectives into natural language instructions. |
| Outcome: | The proposed method outperforms existing methods by 2.4 times in the ASR of three open-source LLMs and GPT-3.5-Turbo. |
False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Emoticons are widely used in digital communication to convey affective intent, yet their safety implications for Large Language Models (LLMs) remain largely unexplored. |
| Approach: | They propose to use ASCII-based emoticons to perform unintended actions in large language models (LLMs) This vulnerability is pervasive, with an average confusion ratio exceeding 38%, and 90% of confused responses yield 'silent failures' authors call on the community to recognize this emerging vulnerability and develop effective mitigation methods to uphold the safety and reliability of human-LLM interactions. |
| Outcome: | The proposed framework exploits emoticon semantic confusion in six LLMs and demonstrates that existing prompt-based mitigations are ineffective. |
Data-centric NLP Backdoor Defense from the Lens of Memorization (2025.findings-naacl)
Copied to clipboard
| Challenge: | Backdoor attacks pose a severe threat to the trustworthiness of DNN-based language models. |
| Approach: | They propose a data-centric defense that extends memorization definitions to fine-grained sentences . they find that duplicated sentence elements are necessary for successful backdoor attacks . |
| Outcome: | The proposed defense outperforms state-of-the-art defenses against backdoor attacks. |
POSTCONDBENCH: Benchmarking Correctness and Completeness in Formal Postcondition Inference (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks emphasize correctness under limited evaluation settings . evaluation of formal specifications is time-consuming, errorprone and requires substantial expertise. |
| Approach: | They propose a multilingual benchmark for evaluating method-level postcondition generation from real-world software. |
| Outcome: | The proposed benchmarks show that evaluation remains a key bottleneck . 420 Python and Java tasks are paired with a high-quality postcondition set . |
Train in Vain: Functionality-Preserving Poisoning to Prevent Unauthorized Use of Code Datasets (2026.findings-acl)
Copied to clipboard
Yuan Xiao, Jiaming Wang, Yuchen Chen, Wei Song, Jun Sun, Shiqing Ma, Yanzhou Mu, Juan Zhai, Chunrong Fang, Jin Song Dong, Zhenyu Chen
| Challenge: | Existing methods for dataset poisoning require full-dataset poison, which breaks code compilability. |
| Approach: | They propose a functionality-preserving poisoning approach that injects short, compilable weak-use fragments into executed code paths. |
| Outcome: | The proposed method contaminates 10% of the dataset while maintaining 100% compilability and functional correctness. |
The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code Generation (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have emerged as the new recommendation engines, surpassing traditional methods in both capability and scope, particularly in code generation. |
| Approach: | They propose to use a dataset to investigate a new type of bias in Large Language Models for code generation, provider bias, to determine whether the model favors specific providers. |
| Outcome: | The proposed model favors services from Google and Amazon, but without explicit directives, and can modify input code to incorporate their preferred providers without user requests. |