Challenge: Recent advances in Multimodal Large Language Models have raised serious safety concerns.
Approach: They propose a method for manipulating the output preference of MLLMs using a preference hijacked image.
Outcome: The proposed method works at inference time and requires no model modifications.

Similar Papers

MLLM-Protector: Ensuring MLLM’s Safety without Hurting Performance (2024.emnlp-main)

Copied to clipboard

Challenge: MLLMs are deployed on limited image-text pairs, which makes them more vulnerable to catastrophic forgetting of their original abilities during safety fine-tuning.
Approach: They propose a plug-and-play strategy that detects harmful visual inputs and transforms harmful ones into harmless ones.
Outcome: The proposed approach mitigates the risks posed by malicious visual inputs without compromising the original performance of MLLMs.
Mitigating Hallucinations in Multi-modal Large Language Models via Image Token Attention-Guided Decoding (2025.naacl-long)

Copied to clipboard

Challenge: Multi-modal large language models (MLLMs) generate plausible but incorrect content, resulting in hallucinations . recent advances in MLLM technology have demonstrated their outstanding performance in a variety of visual tasks, such as object detection.
Approach: They propose a plug-and-play method which leverages MLLMs’ internal representations to mitigate hallucinations by analyzing input and output tokens.
Outcome: The proposed method exploits MLLMs’ internal representations to mitigate hallucinations.
Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection (2026.acl-long)

Copied to clipboard

Challenge: Multimodal large language models are powerful tools for analyzing Internet-scale image data.
Approach: They propose a method to protect images from unauthorized analysis by MLLMs . they embed a perturbation that acts as a visual prompt injection attack on MLMLs if a malicious actor downloads and queries an image .
Outcome: The proposed method protects images from unauthorized analysis by MLLMs . it embeds a perturbation that acts as a visual prompt injection attack on MLMLs if a malicious actor downloads and queries the protected image .
SafeSteer: A Decoding-level Defense Mechanism for Multimodal Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing defense methods rely on fine-tuning or inefficient post-hoc interventions, limiting their ability to address novel attacks.
Approach: They propose a decoding-level defense mechanism that employs a lightweight discriminator to iteratively steer the decoding process toward safety.
Outcome: The proposed method improves safety performance by up to 33.40% without fine-tuning on multiple MLLMs.
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization (2025.findings-acl)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) are known to hallucinate, which limits their practical applications.
Approach: They propose a method that uses three types of preference pairs to target hallucinations from their diverse forms and causes.
Outcome: The proposed method surpasses most state-of-the-art methods and shows potential for further improvements.
Reference Attack: A New Cross-Modal Jailbreaking Attack against Multimodal Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have raised significant safety concerns about generated content, drawing attention from both academia and industry.
Approach: They propose a reference-guided cross-modal jailbreak method that enhances existing prompt-to-image injection attacks by exploiting MLLMs’ semantic reconstruction capabilities.
Outcome: The proposed method achieves an attack success rate of over 93% on leading MLLMs including ChatGPT, Gemini, Claude, and the widely used open-source LLaMA model.
B-APO: Bias-Targeted Adversarial Preference Optimization for Debiasing Multimodal Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing debiasing methods create biased responses by completely removing an entire modality, forming an extreme and static training environment.
Approach: They propose a method to debiase multimodal large language models by masking one modality and then enlarge the margin between clean and adversarial responses.
Outcome: The proposed method achieves superior debiasing performance while maintaining general capabilities.
SURE: Safety Understanding and Reasoning Enhancement for Multimodal Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing multimodal large language models incorporate visual and textual information, but introduces new and complex safety risks.
Approach: They propose a safety reasoning framework that integrates visual modalities into multimodal models to help them resist jailbreak attacks.
Outcome: The proposed framework improves model safety while avoiding over-defense . it is based on a large-scale safety reasoning dataset .
Probing Multimodal Large Language Models for Global and Local Semantic Representations (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have focused on the ability of MLLMs to generate single tokens one by one, while lacking studies about how their representation vectors can encode global multimodal information.
Approach: They propose to use image-caption corpus to train Multimodal Large Language Models (MLLMs) . they find that the topmost layers encode more global semantic information .
Outcome: The proposed models can encode more global semantic information, rather than the topmost layers, and perform better on visual-language entailment tasks.
\mathsf{Con Instruction}: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities (2025.acl-long)

Copied to clipboard

Challenge: Existing attacks communicate instruction through text, accompanied by a toxic image or audio . a novel gray-box attack method generates adversarial images or audio to convey harmful instructions to MLLMs .
Approach: They propose a gray-box attack method that generates adversarial images or audio to convey specific harmful instructions to MLLMs by following non-textual instruction.
Outcome: The proposed method achieves highest success rates on visual and audio-language models . larger models are more susceptible toCon Instruction, compared to their underlying models - the results will be released .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations