Challenge: Existing models that process multiple modalities of data have been used for multimodal tasks, but their advanced capabilities raise privacy concerns.
Approach: They propose a method to modify the model’s internal states associated with PII-related content and to reduce the risk of PI I leakage by modifying the model's internal state.
Outcome: The proposed method achieves on average 93.3% refusal rate for various PII-related tasks with minimal impact on unrelated model performances.

Similar Papers

PII-VisBench: Evaluating Personally Identifiable Information Safety in Vision Language Models Along a Continuum of Visibility (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluations of PII leakage ignore how a subject’s online presence affects privacy alignment.
Approach: They propose a benchmark that evaluates safety through the continuum of online presence by stratifying 200 subjects into four visibility categories: high, medium, low, and zero.
Outcome: The proposed model stratifies 200 subjects into four visibility categories based on the extent and nature of their information available online.
PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing (2026.findings-eacl)

Copied to clipboard

Challenge: Existing defense mechanisms to mitigate PII leakage are limited by existing defenses . a new approach, PATCH, identifies and edits PI I circuits to reduce leakage .
Approach: They propose to use PATCH: Privacy-Aware Targeted Circuit Patching to identify PII leakage circuits in language models to reduce leakage.
Outcome: The proposed approach reduces leakage by up to 65% and can reduce residual leakage to as low as 0.01%.
Unveiling Privacy Risks in Multi-modal Large Language Models: Task-specific Vulnerabilities and Mitigation Challenges (2025.findings-acl)

Copied to clipboard

Challenge: Privacy risks in text-only Large Language Models are well-documented, especially their tendency to memorize and leak sensitive information.
Approach: They propose a dataset to assess privacy risks across multi-modal tasks and scenarios . they demonstrate how models leak sensitive data across various tasks .
Outcome: The proposed model can leak sensitive data embedded in images or stored in memory, exposing privacy risks.
Vision Language Model Helps Private Information De-Identification in Vision Data (2025.findings-acl)

Copied to clipboard

Challenge: Visual Language Models (VLMs) have gained popularity due to their ability to solve imagerelated tasks.
Approach: They propose a framework to enhance privacy awareness of visual language models . they use a specialized instruction-tuning dataset and a tailored training methodology .
Outcome: The proposed framework outperforms existing approaches in handling private information.
Large Language Models Can Be Contextual Privacy Protection Learners (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable linguistic comprehension and generation capability, but when applied to specialized industries, they face challenges such as hallucination, insufficient domain knowledge, and failing to incorporate the latest domain knowledge.
Approach: They propose a paradigm for fine-tuning LLMs that effectively injects domain-specific knowledge while safeguarding inference-time data privacy.
Outcome: The proposed model protects private data while enhancing the model's knowledge.
MLLM-Protector: Ensuring MLLM’s Safety without Hurting Performance (2024.emnlp-main)

Copied to clipboard

Challenge: MLLMs are deployed on limited image-text pairs, which makes them more vulnerable to catastrophic forgetting of their original abilities during safety fine-tuning.
Approach: They propose a plug-and-play strategy that detects harmful visual inputs and transforms harmful ones into harmless ones.
Outcome: The proposed approach mitigates the risks posed by malicious visual inputs without compromising the original performance of MLLMs.
Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations (2026.eacl-long)

Copied to clipboard

Challenge: Existing Med-VLMs are vulnerable to harmful clinical queries . authors propose a novel inference-time defense strategy to mitigate harmful queries based on synthetic clinical demonstrations .
Approach: They propose a novel inference-time defense strategy to mitigate harmful queries . existing Med-VLMs are vulnerable to harmful queries, they argue .
Outcome: The proposed strategy reduces query risk while reducing demonstration budget . existing Med-VLMs are vulnerable to harmful queries, authors argue .
Combating Security and Privacy Issues in the Era of Large Language Models (2024.naacl-tutorials)

Copied to clipboard

Challenge: a tutorial aims to provide a summary of risks and vulnerabilities in large language models . a number of studies have focused on security, privacy and copyright aspects of LLMs .
Approach: This tutorial seeks to provide a systematic summary of risks and vulnerabilities in large language models . authors will discuss security, privacy and copyright aspects of LLMs .
Outcome: This tutorial aims to provide a systematic summary of risks and vulnerabilities in large language models . it will also outline emerging challenges in security, privacy and reliability of LLMs .
PrivLM-Bench: A Multi-level Privacy Evaluation Benchmark for Language Models (2024.acl-long)

Copied to clipboard

Challenge: generative large language models (LLMs) exhibit surprising capability and integrate previous tasks into a unified text generation formulation.
Approach: They propose a privacy evaluation benchmark to quantify the privacy leakage of language models.
Outcome: The proposed benchmark compares PPLMs with different privacy implementations to find out how privacy leakage is handled.
ImgTrojan: Jailbreaking Vision-Language Models with ONE Image (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies on the safety of large language models (LLMs) with human values have focused on the integration of multi-modal user input into these models.
Approach: They propose a method to bypass safety constraints of large language models by using poisoned images instead of original textual captions.
Outcome: The proposed attack bypasses safety constraints of large language models (VLMs) by replacing the original textual captions with malicious jailbreak prompts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations