Papers by Qin Chao
Event Causality Is Key to Computational Story Understanding (2024.naacl-long)
Copied to clipboard
| Challenge: | Cognitive science and symbolic AI research suggest that event causality provides vital information for story understanding. |
| Approach: | They propose a method for event causality identification that leads to material improvements in story understanding. |
| Outcome: | The proposed method improves story understanding on the COPES dataset . it achieves 4.1-10.9% increase on Clip Accuracy and 4.2-13.5% increase on Sentence IoU . |
Knowing-but-Doing: Diagnosing and Defending Role-Play-Driven LLMs Jailbreaks via Moral Disengagement (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used in role-play scenarios, but their safety implications remain under-characterized. |
| Approach: | They propose a diagnostic benchmark for role-play jailbreaks based on Bandura’s Moral Disengagement theory and propose 'MD-Trace' based defense that reduces attack success while maintaining Role Fidelity. |
| Outcome: | The proposed framework improves safety behavior for benign personas while increasing unsafe compliance for malicious ones. |
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches focus on syntactic correctness through synthetic micro-benchmarks or subjective human ratings, despite semantic fidelity and usability. |
| Approach: | They propose a framework that enables effective evaluation of decompilers in reverse engineering workflows . they compare six industrial-strength decompils and six recent LLM-powered approaches . |
| Outcome: | The proposed framework outperforms commercial tools in code understandability despite lower functionality correctness . it shows that it can transform human-centric reverse engineering workflows . |
RU22Fact: Optimizing Evidence for Multilingual Explainable Fact-Checking on Russia-Ukraine Conflict (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods to verify factuality of claims do not provide sufficient evidence for explainable fact-checking systems. |
| Approach: | They propose a method to automatically retrieve and summarize evidence from the Web and a novel multilingual explainable fact-checking dataset on the Russia-Ukraine conflict in 2022. |
| Outcome: | The proposed method can retrieve and summarize evidence from the Web and generate explanations in 16 languages. |
WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning (2025.emnlp-main)
Copied to clipboard
Zhepei Wei, Wenlin Yao, Yao Liu, Weizhi Zhang, Qin Lu, Liang Qiu, Changlong Yu, Puyang Xu, Chao Zhang, Bing Yin, Hyokun Yun, Lihong Li
| Challenge: | Existing work on reinforcement learning has focused on single-turn tasks such as solving math problems. |
| Approach: | They propose a framework that learns directly from online interactions by asynchronously generating diverse trajectories, guided by binary rewards depending on task success. |
| Outcome: | Experiments on the WebArena-Lite benchmark show that the framework outperforms state-of-the-art methods and strong proprietary models. |
Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models (2025.findings-acl)
Copied to clipboard
Qin Liu, Chao Shang, Ling Liu, Nikolaos Pappas, Jie Ma, Neha Anna John, Srikanth Doss, Lluis Marquez, Miguel Ballesteros, Yassine Benajiba
| Challenge: | LLaVA-7B demonstrated a decline in safety alignment ability on multi-modal inputs compared to its LLM backbone. |
| Approach: | They propose a method to recover alignment ability from LLM backbone while preserving functional capabilities of VLMs. |
| Outcome: | The proposed framework recovers alignment ability that is inherent in the LLM backbone with minimal impact on fluency and linguistic capabilities of pre-trained VLMs. |
UrbanLLM: Autonomous Urban Activity Planning and Management with Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | UrbanLLM is a fine-tuned large language model designed to tackle diverse urban problems. |
| Approach: | They propose a fine-tuned large language model to tackle diverse urban problems . UrbanLLM decomposes urban-related queries into manageable sub-tasks . |
| Outcome: | The proposed model outperforms existing models in urban planning and management tasks. |
TransLLM: A Unified Multi-Task Large Language Model for Urban Transportation via Learnable Prompting (2026.acl-long)
Copied to clipboard
| Challenge: | Existing models lack generalization capabilities and lack structured spatiotemporal data. |
| Approach: | They propose a unified multi-task framework that synergizes spatiotemporal encoding with LLM reasoning through learnable prompt composition. |
| Outcome: | The proposed framework outperforms baseline models on seven datasets and three tasks on supervised and zero-shot settings with excellent generalization and robustness. |
GenDis: Generative-Discriminative Dual-View Co-Training for Generalized Category Discovery (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods rely on one-hot discriminative supervision, leading to overfitting on seen classes and poor generalization to unseen ones. |
| Approach: | They propose a Generative–Discriminative Dual-View Co-Training framework that unifies discriminative classification and semantic label generation within an LLM. |
| Outcome: | The proposed framework outperforms existing methods on five benchmarks on the generalized category discovery (GCD) task. |
PLaD: Preference-based Large Language Model Distillation with Pseudo-Preference Pairs (2024.findings-acl)
Copied to clipboard
Rongzhi Zhang, Jiaming Shen, Tianqi Liu, Haorui Wang, Zhen Qin, Feng Han, Jialu Liu, Simon Baumgartner, Michael Bendersky, Chao Zhang
| Challenge: | Knowledge distillation (KD) is a technique for transferring expertise from large teacher models to compact student models with reduced memory footprints and inference costs. |
| Approach: | They propose to transfer knowledge from large teacher models to compact student models by exploiting teacher-student capacity discrepancies to generate pseudo-preference pairs where teacher outputs are preferred over student outputs. |
| Outcome: | The proposed framework exploits teacher-student capacity discrepancy to generate pseudo-preference pairs where teacher outputs are preferred over student outputs. |
MDC-Bench: A Multidisciplinary Causal Benchmark Based on Causal Structures for Evaluating Large Language Models (2026.findings-acl)
Copied to clipboard
Peng Wang, Yuxiong Yan, Xiao Ding, Kai Xiong, Bibo Cai, Chao Peng, Yutai Hou, Dandan Tu, Bing Qin, Ting Liu
| Challenge: | Existing causal datasets focus on the commonsense domain, but LLMs perform poorly when answering complex questions. |
| Approach: | They propose a multidisciplinary causal evaluation benchmark to assess LLMs' knowledge and skills. |
| Outcome: | The proposed model improves in domain specialization, structural diversity, and task complexity. |
Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods to learn internal world models rely on one-step supervision . however, standard MTP suffers from structural hallucinations . |
| Approach: | They propose a method which anchors predictions to ground-truth hidden state trajectories. |
| Outcome: | The proposed method bridges the gap between discrete tokens and continuous state representations, reducing structural hallucinations, and improving robustness to perturbations. |
Zero-to-Strong Generalization: Eliciting Strong Capabilities of Large Language Models Iteratively without Gold Labels (2025.coling-main)
Copied to clipboard
| Challenge: | Pre-trained language models have demonstrated remarkable performance through supervised fine-tuning or in-context learning using gold labels. |
| Approach: | They propose a new paradigm termed zero-to-strong generalization that prompts LLMs to annotate unlabeled data and retain high-quality labels by filtering. |
| Outcome: | The proposed framework outperforms pre-trained language models on extensive classification and reasoning tasks on multiple model sizes. |
PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing research focuses on character-level settings and static evaluation formats fail to capture the complexity of everyday social interactions. |
| Approach: | They propose a dynamic simulation framework for evaluating and improving persona-level role-playing in large language models (LLMs). |
| Outcome: | The proposed framework leverages user-generated social content to construct a nuanced persona bank and elicits multi-turn, context-rich interactions within simulated social environments. |
Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Semantic phrases (SP) are lexical combinations whose meanings or usages may not be fully derived from their individual components. |
| Approach: | They propose to consolidate existing multiword expression resources into a unified testbed to assess language models in semantic phrase processing tasks. |
| Outcome: | The evaluation suite covers idiomatic expressions, noun compounds, and verbal constructions. |
MessToClean: Evidence-Grounded Structure-Preserving Reconstruction for Real-World Degraded Exam Paper Images (2026.acl-long)
Copied to clipboard
Jiayi Tuo, Cheng Tang, Zihan Wang, Chenyue Zhou, Yao Li, Yanbiao Ma, Chao Wang, Wei Dai, Mingxuan Wang, Shitong Qin, Ziwei Zhao
| Challenge: | Existing Multimodal Large Language Models (MLLMs) fail under RDEI, leading to disrupted structure and evidence-unsupported hallucinations. |
| Approach: | They propose a backbone-agnostic, evidence-driven pipeline that treats off-the-shelf MLLMs as interchangeable components to improve stem consistency and figure consistency. |
| Outcome: | The proposed pipeline improves stem consistency by 1.01-3.18%, figure consistency by 0.50-49.16%, and refusal F1 by 1.06-10.88% across question types. |
Explanation-aware Soft Ensemble Empowers Large Language Model In-context Learning (2024.acl-long)
Copied to clipboard
Yue Yu, Jiaming Shen, Tianqi Liu, Zhen Qin, Jing Nathan Yan, Jialu Liu, Chao Zhang, Michael Bendersky
| Challenge: | Recent advances in natural language processing (NLP) have witnessed the remarkable capabilities of Large Language Models (LLMs). |
| Approach: | They propose an Explanation-Aware Soft Ensemble framework to empower in-context learning with Large language models. |
| Outcome: | The proposed framework can be used to enhance in-context learning on seven natural language understanding tasks and four varying-size LLMs. |