Papers by Ziming Li
Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches focus on improving attack success rates while overlooking the need for comprehensive test case coverage. |
| Approach: | They propose a top-down approach to automated red teaming that scales up the diversity of test cases using an extensible, fine-grained risk taxonomy. |
| Outcome: | The proposed approach scales up the diversity of test cases using a top-down approach based on an extensible, fine-grained risk taxonomy and leverages reinforcement learning techniques to facilitate multi-turn adversarial probing in a human-like manner. |
Know Your Place: Diagnosing Implicit Social Adaptation Failures in Chinese Large Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing studies suggest that failures of large language models in social contexts are not due to limited linguistic competence, but to inappropriate recognition. |
| Approach: | They propose a framework that decomposes social adaptation into three orthogonal dimensions and conduct controlled comparisons across multiple Chinese LLMs under implicit and explicit conditions. |
| Outcome: | The proposed framework decomposes social adaptation into three orthogonal dimensions and conducts controlled comparisons across multiple Chinese LLMs under implicit and explicit conditions. |
Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context Learning (2025.acl-long)
Copied to clipboard
Shaobo Wang, Xiangqi Jin, Ziming Wang, Jize Wang, Jiajun Zhang, Kaixin Li, Zichen Wen, Zhong Li, Conghui He, Xuming Hu, Linfeng Zhang
| Challenge: | Using fine-tuning on task-specific data is essential for large language models to be effective in specialized tasks. |
| Approach: | They propose a method that leverages few-shot in-context learning with the model to be fine-tuned. |
| Outcome: | The proposed method outperforms existing methods with a 3.1-point improvement and a 7.4 speedup on the Llama-3-8B-Instruct model using just 10% of the dataset. |
QAP: A Quantum-Inspired Adaptive-Priority-Learning Model for Multimodal Emotion Recognition (2023.findings-acl)
Copied to clipboard
| Challenge: | Experimental results show that multimodal emotion recognition is a state-of-the-art technique . textual, visual and acoustic modalities are involved in multimodal video emotion recognition . |
| Approach: | They propose a quantum-inspired adaptive-priority-learning model to address the challenges . they use quantum state to model modal features and Q-attention to integrate three modalities . |
| Outcome: | Experimental results show that QAP improves on previous models. |
MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning (2024.findings-acl)
Copied to clipboard
Xiangru Tang, Anni Zou, Zhuosheng Zhang, Ziming Li, Yilun Zhao, Xingyao Zhang, Arman Cohan, Mark Gerstein
| Challenge: | Large language models face unique challenges such as domain-specific terminologies and reasoning over specialized knowledge. |
| Approach: | They propose a multi-disciplinary collaboration framework that leverages LLM-based agents in a role-playing setting. |
| Outcome: | The proposed framework excels at mining and harnessing medical expertise within LLMs, as well as extending its reasoning abilities. |
BiBL: AMR Parsing and Generation with Bidirectional Bayesian Learning (2022.coling-1)
Copied to clipboard
| Challenge: | Existing approaches to AMR focus on one-side improvements despite the duality of the two tasks . instead, we propose data-efficient Bidirectional Bayesian learning (BiBL) to facilitate bidirectional information transition. |
| Approach: | They propose a data-efficient bidirectional Bayesian learning approach to facilitate bidirectional information transition by adopting a single-stage multitasking strategy. |
| Outcome: | The proposed model outperforms existing models on benchmark datasets without extra training data. |
Gated Mechanism Enhanced Multi-Task Learning for Dialog Routing (2022.coling-1)
Copied to clipboard
| Challenge: | Existing methods for dialog routing are mostly heuristic and cannot achieve high-quality performance. |
| Approach: | They propose a multi-task learning framework with a dialog encoder and two tailored gated mechanism modules to solve this problem. |
| Outcome: | The proposed model can play the role of hierarchical information filtering and is non-invasive to existing dialog systems. |
See the World, Discover Knowledge: A Chinese Factuality Evaluation for Large Vision Language Models (2025.findings-acl)
Copied to clipboard
Jihao Gu, Yingyao Wang, Pi Bu, Chen Wang, Ziming Wang, Tengtao Song, Donglai Wei, Jiale Yuan, Yingxiu Zhao, Yancheng He, Shilong Li, Jiaheng Liu, Meng Cao, Jun Song, Yingshui Tan, Xiang Li, Wenbo Su, Xiaoyong Zhu, Bo Zheng
| Challenge: | Existing models for large vision language models do not fully reflect their knowledge capacity and reliability, resulting in erroneous outputs that do not align with the image content or provide answers lacking knowledge evidence. |
| Approach: | They propose a Chinese-based benchmark for visual factuality across 8 major topics and 56 subtopics and a multi-hop question construction. |
| Outcome: | The proposed model decouples visual factuality into two parts: seeing the world and discovering knowledge. |
Explainable Quantum Program Repair with Verifiable Proof Traces (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to program repair provide only post-hoc, non-verifiable explanations that are not executable or verifiably. |
| Approach: | They propose a framework that couples repair generation with machine-checkable executable explanations for quantum programs where correctness hinges on subtle semantic properties such as circuit equivalence and fidelity preservation. |
| Outcome: | Experiments on QASMBench with mutation-generated quantum program bugs show that the proposed framework improves both semantic precision and explanation faithfulness over baselines that rely on unconstrained or purely natural-language explanations. |
Rethinking Supervised Learning and Reinforcement Learning in Task-Oriented Dialogue Systems (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Dialogue policy learning for task-oriented dialogue systems has enjoyed great progress through using reinforcement learning methods. |
| Approach: | They propose a dialogue action decoder and a simulator-free adversarial learning method to improve dialogue agent performance without using reinforcement learning. |
| Outcome: | The proposed methods achieve more stable and higher performance with fewer efforts, such as the domain knowledge required to design a user simulator and the intractable parameter tuning in reinforcement learning. |
Towards Objective Fine-tuning: How LLMs’ Prior Knowledge Causes Potential Poor Calibration? (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have enabled powerful domain-specific applications through supervised fine-tuning. |
| Approach: | They propose a cognition-aware framework that applies targeted learning strategies according to the model’s prior knowledge to improve calibration. |
| Outcome: | The proposed framework significantly improves calibration while maintaining performance, achieving an average 57% reduction in ECE compared to standard fine-tuning in Llama3-8B. |
TrojanSQL: SQL Injection against Natural Language Interface to Database (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on text-to-SQL systems have not investigated its security aspects . however, how to implement such attacks remains an open question. |
| Approach: | They propose a backdoor-based SQL injection framework for text-to-SQL systems that uses boolean-based injection and union-based injecting techniques to exploit SQL injection vulnerabilities. |
| Outcome: | The proposed framework can produce harmful SQL statements invalidating user queries or compromise sensitive information about the database. |
Investigating Crowdsourcing Protocols for Evaluating the Factual Consistency of Summaries (2022.naacl-main)
Copied to clipboard
Xiangru Tang, Alexander Fabbri, Haoran Li, Ziming Mao, Griffin Adams, Borui Wang, Asli Celikyilmaz, Yashar Mehdad, Dragomir Radev
| Challenge: | Existing pre-trained summarization models produce text that is factually inconsistent with the input. |
| Approach: | They present a scale-based scale for Likert rating and a scoring algorithm for Best-Worst Scaling to improve crowdsourcing reliability. |
| Outcome: | The proposed model is more reliable than existing models on two news summarization datasets. |
FanChuan: A Multilingual and Graph-Structured Benchmark For Parody Detection and Analysis (2025.findings-acl)
Copied to clipboard
Yilun Zheng, Sha Li, Fangkun Wu, Yang Ziyi, Lin Hongchao, Zhichao Hu, Cai Xinjun, Ziming Wang, Jinxuan Chen, Sitao Luan, Jiahao Xu, Lihui Chen
| Challenge: | Parody is an emerging phenomenon on social media, where individuals imitate a role or position opposite to their own . limited available data and deficient diversity in current datasets hinder study of parody . |
| Approach: | They build a dataset of parody users and annotated comments from both English and Chinese corpora to test parody detection and comment sentiment analysis. |
| Outcome: | The proposed datasets provide richer contextual information, which is lacking in existing datasets. |
FG-PRM: Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to detect hallucinations in large language models lack nuanced understanding of their types and manifestations. |
| Approach: | They propose a taxonomy that categorizes hallucinations into six types . they propose an augmented model to detect and mitigate hallucinosity in a fine-grained manner . |
| Outcome: | The proposed model detects and mitigates hallucinations in a fine-grained manner . it significantly boosts the performance of LLMs on GSM8K and MATH benchmarks. |
LMR-BENCH: Evaluating LLM Agent’s Ability on Reproducing Language Modeling Research (2025.emnlp-main)
Copied to clipboard
Shuo Yan, Ruochen Li, Ziming Luo, Zimu Wang, Daoyang Li, Liqiang Jing, Kaiyu He, Peilin Wu, Juntong Ni, George Michalopoulos, Yue Zhang, Ziyang Zhang, Mian Zhang, Zhiyu Chen, Xinya Du
| Challenge: | Large language model (LLM) agents have demonstrated remarkable potential in advancing scientific discovery, but their capability in reproducing code from research papers remains underexplored. |
| Approach: | They propose to evaluate LLM agents' ability to reproduce scientific research papers by analyzing code reproduction tasks from 23 research papers published in top-tier NLP venues. |
| Outcome: | The proposed benchmark systematically evaluates the capability of large language model (LLM) agents on code reproduction from Language Modeling Research. |
Debiasing LLMs by Masking Unfairness-Driving Attention Heads (2026.findings-acl)
Copied to clipboard
Tingxu Han, Wei Song, Ziqi Ding, Ziming Li, Chunrong Fang, Yuekang Li, Dongfang Liu, Zhenyu Chen, Zhenting Wang
| Challenge: | Existing work probes when biased outputs appear, but gives little insight into the mechanisms that generate them, leaving existing mitigations largely fragile. |
| Approach: | They propose a lightweight debiasing framework that detects bias heads and selectively masks only those heads that activate under DA and CoT. |
| Outcome: | The proposed framework reduces unfairness by 391.9%- 534.5% in both one- and two-turn dialogues. |
Making Task-Oriented Dialogue Datasets More Natural by Synthetically Generating Indirect User Requests (2025.coling-main)
Copied to clipboard
| Challenge: | Existing task-oriented dialogue benchmarks lack sufficient examples of complex discourse phenomena such as indirectness. |
| Approach: | They propose a set of linguistic criteria and an LLM-based pipeline for generating realistic IURs to test natural language understanding and dialogue state tracking models before deployment in a new domain. |
| Outcome: | The proposed model can handle indirect user requests (IURs) but lacks examples of complex discourse phenomena such as indirectness. |
Guided Dialogue Policy Learning without Adversarial Learning in the Loop (2020.findings-emnlp)
Copied to clipboard
Ziming Li, Sungjin Lee, Baolin Peng, Jinchao Li, Julia Kiseleva, Maarten de Rijke, Shahin Shayandeh, Jianfeng Gao
| Challenge: | Reinforcement learning methods suffer from sparse and unstable reward signals . alternating training of dialogue agent and reward model can get stuck in local optima . |
| Approach: | They propose to decompose adversarial training into two steps to improve dialogue policy learning. |
| Outcome: | The proposed method achieves remarkable task success rate using both on-policy and off-poly reinforcement learning methods. |
Born Pragmatic, Trained to Hallucinate? Quantifying the Origins of Contextual Bias in LLMs via the PaCE Benchmark (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models excel at capturing communicative intent, but they have a side effect: pragmatic hallucination. |
| Approach: | They propose a benchmark to quantify the impact of pragmatic hallucination on large language models . they propose RLHF and SFT to induce a strong tendency for pragmatic over-attribution . |
| Outcome: | The proposed model outperforms existing models in predicting pragmatic hallucinations . the evaluations show that current alignment paradigms lack precise control over pragmatic boundaries . |
A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code (2026.findings-acl)
Copied to clipboard
Keke Lian, Wang Bin, Lei Zhang, Libo Chen, Junjie Wang, Ziming Zhao, Yujiu Yang, Miaoqian Lin, Haotong Duan, Haoran Zhao, Shuang Liao, Mingda Guo, Quan Jiazheng, Yilu Zhong, Chenhao He, Chen Zichuan, Jie Wu, Haoling Li, Zhaoxuan Li, Jiongchi Yu, Hui LI, Dong Zhang
| Challenge: | Existing security evaluation benchmarks lack relevance to real-world AI programming tasks . current LLMs struggle with secure coding, research shows . |
| Approach: | They propose a repository-level evaluation benchmark to assess security of AI-generated code. |
| Outcome: | The proposed framework mirrors real-world AI programming tasks and offers valuable insights into the state of AI code generation. |
AMOA: Global Acoustic Feature Enhanced Modal-Order-Aware Network for Multimodal Sentiment Analysis (2022.coling-1)
Copied to clipboard
| Challenge: | Existing methods treat three modal features equally, without distinguishing the importance of different modalities. Existing models split the video into frames, leading to missing the global acoustic information. |
| Approach: | They propose a global Acoustic feature enhanced Modal-Order-Aware network to address these problems. |
| Outcome: | The proposed model outperforms state-of-the-art models on two public datasets. |