Papers by Yulin Wu
Defense Against Prompt Injection Attack by Leveraging Attack Techniques (2025.acl-long)
Copied to clipboard
| Challenge: | Recent attacks leverage LLMs’ instruction-following abilities and their inabilities to distinguish instructions injected in the data content. |
| Approach: | They invert the intention of prompt injection methods to develop novel defense methods based on previous training-free attack methods by repeating the attack process with the original input instruction rather than the injected instruction. |
| Outcome: | The proposed methods outperform existing defense approaches, achieving state-of-the-art results. |
EvoMD-LLM: Learning the Language of Species Evolution in Reactive Molecular Dynamics (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing models operate on static molecular representations or rely on external tools for reasoning. |
| Approach: | They propose a framework that reformulates species-level molecular dynamics as a symbolic temporal language modeling problem. |
| Outcome: | The proposed model outperforms neural networks and language-based baselines on multiple temporal prediction tasks and generates plausible interpretations of reaction dynamics. |
MPO: Multilingual Safety Alignment via Reward Gap Optimization (2025.acl-long)
Copied to clipboard
Weixiang Zhao, Yulin Hu, Yang Deng, Tongtong Wu, Wenxuan Zhang, Jiahe Guo, An Zhang, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu
| Challenge: | Existing preference learning methods for safety alignment are monolingual and struggle with noisy multilingual data. |
| Approach: | They propose a multilingual reward gaP optimization approach that leverages the well-aligned safety capabilities of the dominant language to improve safety alignment across multiple languages. |
| Outcome: | Extensive experiments on three LLMs, LLaMA-3.1, Gemma-2 and Qwen2.5, validate MPO’s efficacy in multilingual safety alignment without degrading general multilingual utility. |
Do Current Video LLMs Have Strong OCR Abilities? A Preliminary Study (2025.coling-main)
Copied to clipboard
| Challenge: | a new benchmark evaluates video-based optical character recognition (Video OCR) performance of multi-modal models in videos . the benchmark aims to improve video LLMs' ability to extract text from video content . previous benchmarks have focused on video QA, but not video-related QA. |
| Approach: | They propose to evaluate the video OCR performance of multi-modal models in videos . they use a semi-automated approach that integrates the OCR ability of image LLMs with manual refinement . |
| Outcome: | The proposed benchmark includes 1,028 videos and 2,961 question-answer pairs . it integrates the OCR ability of image LLMs with manual refinement . |
When Retriever-Reader Meets Scenario-Based Multiple-Choice Questions (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Scenario-based question answering (SQA) requires retrieving and reading paragraphs from a large corpus to answer a question contextualized by a long scenario description. |
| Approach: | They propose a model where the retriever is implicitly supervised only using QA labels via a novel word weighting mechanism. |
| Outcome: | The proposed model outperforms strong baselines on multiple-choice questions in three datasets. |
LaERC-S: Improving LLM-based Emotion Recognition in Conversation with Speaker Characteristics (2025.coling-main)
Copied to clipboard
| Challenge: | Emotion recognition in conversation (ERC) is a task of discerning human emotions for each utterance within a conversation. |
| Approach: | They propose a framework that uses large language models to analyze speaker characteristics . they use two-stage learning to make the models reason speaker characteristics and track emotion of the speaker . |
| Outcome: | The proposed framework outperforms existing methods on three benchmark datasets. |