Papers by Jiachen Ma
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | AVLM integrates full-face visual cues into a pre-trained expressive speech model. |
| Approach: | They propose an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. |
| Outcome: | The proposed model incorporates full-face visual cues into a pre-trained expressive speech model. |
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Text-to-image (T2I) models can be used to generate harmful content such as sexually explicit, unfaithful, and misleading or Not-Safe-for-Work (NSFW) images. |
| Approach: | They propose a more practical and universal attack that does not require the presence of a target model. |
| Outcome: | The proposed attack bypasses both text and image safety checkers while preserving high semantic alignment with the target prompt. |