Papers with Gemini-2.5-pro
Audio-Aware Large Language Models as Judges for Speaking Styles (2025.findings-emnlp)
Copied to clipboard
Cheng-Han Chiang, Xiaofei Wang, Chung-Ching Lin, Kevin Lin, Linjie Li, Radu Kopetz, Yao Qian, Zhendong Wang, Zhengyuan Yang, Hung-yi Lee, Lijuan Wang
| Challenge: | Audio-aware large language models (ALLMs) can understand textual and non-textual information in the audio input. |
| Approach: | They use audio-aware large language models (ALLMs) to evaluate the speaking styles of SLMs on two tasks: voice style instruction following and role-playing. |
| Outcome: | The proposed models can understand the textual and non-textual information in the audio input and can be used as a judge to assess the speaking styles of SLMs. |
DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs (2026.acl-long)
Copied to clipboard
| Challenge: | Existing jailbreak methods only use a single image, restricting the attack space . Existing frameworks only use single image to distribute harmful requests across multiple images . |
| Approach: | They propose a compositional jailbreak framework that leverages Distributed instruction, Multimodal evidence and a Number chain task to fully enhance the jailbreak performance. |
| Outcome: | The proposed framework achieves attack success rates of over 90% on GPT-4o, Gemini-2.5-pro and Claude Sonnet 4 . |
The “Knowledge–Behavior Gap” in Cultural Taboo Safety of Large Language Models (2026.acl-long)
Copied to clipboard
Ying He, Sihang Jiang, Xingzhou Chen, Zhouhong Gu, Yiwei Gu, Minggui HE, Shimin Tao, null Mahongxia, Yanghua Xiao
| Challenge: | Existing cultural benchmarks assess cultural knowledge or values biases, but ignore cultural taboos. |
| Approach: | They propose a benchmark to evaluate and improve the cultural taboo safety of large language models. |
| Outcome: | The proposed benchmark spans 77 countries and regions, and includes over 2,020 taboos. |