Papers by Zhouxing Shi
PII-VisBench: Evaluating Personally Identifiable Information Safety in Vision Language Models Along a Continuum of Visibility (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluations of PII leakage ignore how a subject’s online presence affects privacy alignment. |
| Approach: | They propose a benchmark that evaluates safety through the continuum of online presence by stratifying 200 subjects into four visibility categories: high, medium, low, and zero. |
| Outcome: | The proposed model stratifies 200 subjects into four visibility categories based on the extent and nature of their information available online. |
Defending LLMs against Jailbreaking Attacks via Backtranslation (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent advancement in large language models (LLMs) has shown their extensive applications and transformative potential to reshape people's lives. |
| Approach: | They propose a method which uses backtranslation to infer an input prompt from an input input prompt and then run it again on the backtranslated prompt. |
| Outcome: | The proposed method outperforms baselines and has little impact on the generation quality for benign input prompts. |
From Individual to Common: An Early Exploration of Consensus in Non-verifiable Data for Balanced Preference Optimization (2026.acl-long)
Copied to clipboard
| Challenge: | Reinforcement Learning with Verifiable Rewards (RLVR) has demonstrated remarkable effectiveness in boosting the objective performance of Large Language Models (LLMs). |
| Approach: | They propose a dataset where response pairs differ only by subtle nuances and a model with a non-verifiable dataset. |
| Outcome: | The proposed model outperforms models trained on data with explicit quality gaps while maintaining objective capabilities. |
Robustness to Modification with Shared Words in Paraphrase Identification (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Paraphrase identification models have been shown to be vulnerable and lack robustness in tasks such as text classification and natural language inference. |
| Approach: | They propose to modify an example such that a target model makes a wrong prediction by using beam search constrained by heuristic rules and a BERT-masked language model to generate substitution words compatible with the context. |
| Outcome: | The proposed model performance drops dramatically on modified examples, revealing the robustness issue. |
On the Sensitivity and Stability of Model Interpretations in NLP (2022.acl-long)
Copied to clipboard
| Challenge: | Recent years have witnessed the emergence of post-hoc interpretations that aim to uncover how NLP models make predictions. |
| Approach: | They propose two new criteria that provide complementary notions of faithfulness to removal-based criteria. |
| Outcome: | The proposed methods overcome limitations of gradient-based methods on removal-based criteria and overcome limitations in the proposed methods. |