Papers by Seungjong Sun
Personality Vector: Modulating Personality of Large Language Models by Model Merging (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to induce personality in large language models (LLMs) fail to capture the continuous nature of human traits. |
| Approach: | They propose a method for personality modulation in large language models by model merging by subtracting weights of pre-trained models from those of fine-tuned models. |
| Outcome: | The proposed method allows LLMs to exhibit desired personality traits without additional training. |
Kiss up, Kick down: Exploring Behavioral Changes in Multi-modal Large Language Models with Assigned Visual Personas (2024.emnlp-main)
Copied to clipboard
Seungjong Sun, Eungu Lee, Seo Baek, Seunghyun Hwang, Wonbyung Lee, Dongyan Nan, Bernard Jansen, Jang Kim
| Challenge: | Large language models (LLMs) exhibit a high degree of alignment with human behavior based on their robust capabilities for natural language understanding and generation. |
| Approach: | They developed a dataset of 5K fictional avatar images for assignment as visual personas to large language models (LLMs) and analyzed their negotiation behaviors based on the visual traits depicted in these images. |
| Outcome: | The proposed model exhibited aggressive negotiation behaviors when the opponent’s image appeared less aggressive than their own, and less aggressive negotiation behavior when the opposing image appeared more aggressive. |
Jailbreaking Multimodal Large Language Models using Multi-Clip Video (2026.acl-long)
Copied to clipboard
| Challenge: | Existing studies show that video inputs can bypass safety alignment, yet it remains unclear which properties of video input induce this vulnerability. |
| Approach: | They propose a simple image-based defense that mitigates the vulnerability of MLLMs by analyzing video inputs. |
| Outcome: | The proposed defense leverages the relative robustness of the image modality. |