Papers by Hyunmin Jun
Jailbreaking Multimodal Large Language Models using Multi-Clip Video (2026.acl-long)
Copied to clipboard
| Challenge: | Existing studies show that video inputs can bypass safety alignment, yet it remains unclear which properties of video input induce this vulnerability. |
| Approach: | They propose a simple image-based defense that mitigates the vulnerability of MLLMs by analyzing video inputs. |
| Outcome: | The proposed defense leverages the relative robustness of the image modality. |