Papers by Huan Zheng
Defending Against Social Engineering Attacks in the Age of LLMs (2024.emnlp-main)
Copied to clipboard
Lin Ai, Tharindu Kumarage, Amrita Bhattacharjee, Zizhou Liu, Zheng Hui, Michael Davinroy, James Cook, Laura Cassani, Kirill Trapeznikov, Matthias Kirchner, Arslan Basharat, Anthony Hoogs, Joshua Garland, Huan Liu, Julia Hirschberg
| Challenge: | Existing research has developed frameworks to understand human-to-human CSE attacks. |
| Approach: | They propose a modular defense pipeline that improves detection at both the message and conversation levels. |
| Outcome: | The proposed model can be exploited to facilitate chat-based social engineering attacks and generate high-quality CSE content, but their detection capabilities are suboptimal, leading to increased operational costs for defense. |
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark (2025.acl-long)
Copied to clipboard
Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Shengbang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, Yu Su, Wenhu Chen, Graham Neubig
| Challenge: | Recent advances in multimodal large language models have led to progress in tackling complex reasoning tasks that combine textual and visual information. |
| Approach: | They introduce a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark. |
| Outcome: | The proposed model performs lower on MMMU-Pro than on the previous benchmark, ranging from 16.8% to 26.9%. |
Sibyl: Empowering Empathetic Dialogue Generation in Large Language Models via Sensible and Visionary Commonsense Inference (2025.coling-main)
Copied to clipboard
Lanrui Wang, Jiangnan Li, Chenxu Yang, Zheng Lin, Hongyin Tang, Huan Liu, Yanan Cao, Jingang Wang, Weiping Wang
| Challenge: | Recent studies have focused on integrating commonsense knowledge into chatbots to enhance their ability to understand and generate dialogue responses. |
| Approach: | They propose a framework that integrates commonsense knowledge into chatbots to enable them to elicit more empathetic responses. |
| Outcome: | The proposed framework enables LLMs to uncover the implicit requirements of the conversation, aiming to elicit more empathetic responses. |
WebOlympus: An Open Platform for Web Agents on Live Websites (2024.emnlp-demo)
Copied to clipboard
| Challenge: | Web agents are emerging as powerful tools for automating tasks in cyberspace . however, there is a lack of standardized and user-friendly tools for research and development . |
| Approach: | They propose an open platform for web agents operating on live websites with a Chrome extension and a safety monitor module to ensure their trustworthiness. |
| Outcome: | WebOlympus is an open platform for web agents operating on live websites. |
Multimodal Large Language Models for Multi-Subject In-Context Image Generation (2026.acl-long)
Copied to clipboard
| Challenge: | Recent advances in text-to-image (T2I) generation have enabled visually coherent image synthesis from descriptions, but generating images containing multiple given subjects remains challenging. |
| Approach: | They propose a model that enables automatic and scalable data generation without manual annotations to overcome the data scarcity. |
| Outcome: | The proposed model overcomes the data scarcity and lacks manual annotations. |
Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision (2026.acl-long)
Copied to clipboard
| Challenge: | Existing models with static prompts, rules, or reward models are constrained by static supervision, which often fails to shape the underlying reasoning process, leading to brittle generalization and performance saturation in complex decision-making tasks. |
| Approach: | They propose a principle-centric learning framework that treats reasoning principles as explicit, language-based supervision signals that can be generated, evaluated, and iteratively evolved. |
| Outcome: | The proposed framework treats reasoning principles as explicit, language-based supervision signals that can be generated, evaluated, and iteratively evolved. |