Papers by Yiqun Wang
STICKERCONV: Generating Multimodal Empathetic Responses from Scratch (2024.acl-long)
Copied to clipboard
Yiqun Zhang, Fanheng Kong, Peidong Wang, Shuang Sun, SWangLing SWangLing, Shi Feng, Daling Wang, Yifei Zhang, Kaisong Song
| Challenge: | Prior studies on stickers focused on sentiment analysis and recommendation systems, overlooking their vast potential in empathetic response generation. |
| Approach: | They propose a multimodal empathetic dialogue dataset, STICKERCONV, which simulates human behavior with stickers, and propose evaluative metrics based on LLM. |
| Outcome: | The proposed framework generates contextually relevant and emotionally resonant multimodal empathetic responses, contributing to the advancement of more nuanced and engaging e-dialog systems. |
SelfRACG: Enabling LLMs to Self-Express and Retrieve for Code Generation (2025.emnlp-main)
Copied to clipboard
Qian Dong, Jia Chen, Qingyao Ai, Hongning Wang, Haitao Li, null Yiwu, Yao Hu, Yiqun Liu, Shaoping Ma
| Challenge: | Existing retrieval-augmented code generation methods fail to accurately fetch the knowledge required for code generation for consecutive code fragments. |
| Approach: | They propose a paradigm that enables large language models to Self-express their information needs to enhance retrieval-augmented code generation methods. |
| Outcome: | Experiments show that SelfRACG can retrieve external knowledge that better aligns with the LLM’s own information needs, resulting in superior generation performance compared to vanilla RACG. |
Importance of Synthesizing High-quality Data for Text-to-SQL Parsing (2023.findings-acl)
Copied to clipboard
Yiqun Hu, Yiyun Zhao, Jiarong Jiang, Wuwei Lan, Henghui Zhu, Anuj Chauhan, Alexander Hanbo Li, Lin Pan, Jun Wang, Chung-Wei Hang, Sheng Zhang, Jiang Guo, Mingwen Dong, Joseph Lilien, Patrick Ng, Zhiguo Wang, Vittorio Castelli, Bing Xiang
| Challenge: | Existing text-to-SQL parsers lack the data to perform well with augmented synthetic data. |
| Approach: | They propose a framework that imposes strong typing constraints and incorporates key relationships from schema. |
| Outcome: | The proposed framework improves on the high-quality synthesized SQL and natural language question (NLQ) models have significant accuracy boosts and achieve new state-of-the-art performance on spider. |
If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMs (2026.acl-long)
Copied to clipboard
Siqi Fan, Xiusheng Huang, Yiqun Yao, Xuezhi Fang, Kang Liu, Peng Han, Shuo Shang, Aixin Sun, Yequan Wang
| Challenge: | Existing benchmarks for large language models (LLMs) fail to capture these dynamics, focusing on static, open-ended evaluations. |
| Approach: | They propose a benchmark to assess lifelong learning in large language models . they use two episodic datasets rich in narrative structure and character interactions . |
| Outcome: | Experiments on LLMs show that non-parametric methods outperform parametric ones in managing stateful learning. |
Tracing and Dissecting How LLMs Recall Factual Knowledge for Real World Questions (2025.acl-long)
Copied to clipboard
| Challenge: | Recent advances in large language models have shown promising ability to perform commonsense reasoning. |
| Approach: | They propose a two-dimensional analysis framework that incorporates token back-tracing and token decoding to uncover how LLMs conduct factual knowledge recall. |
| Outcome: | The proposed framework shows that LLMs lack relevant knowledge but struggle to select the most accurate information based on context during the retrieval and rerank phase. |
DeepMed: Building a Medical DeepResearch Agent via Multi-hop Med-Search Data and Turn-Controlled Agentic Training & Inference (2026.findings-acl)
Copied to clipboard
Zihan Wang, Hao Wang, Shi Feng, Xiaocui Yang, Daling Wang, Yiqun Zhang, Jinghao Lin, Xiaozhong Ji, Haihua Yang
| Challenge: | Medical reasoning models are constrained by parametric knowledge and can induce hallucinations and spurious attributions. |
| Approach: | They propose a model that uses a multi-hop med-search QA synthesis method to apply the DR paradigm in medical contexts. |
| Outcome: | The proposed model outperforms larger medical reasoning models on medical benchmarks. |
Why Do More Experts Fail? A Theoretical Analysis of Model Merging (2026.acl-long)
Copied to clipboard
Zijing Wang, Xingle Xu, YongKang Liu, Yiqun Zhang, Peiqin Lin, Shi Feng, Daling Wang, Xiaocui Yang, Hinrich Schuetze
| Challenge: | Existing methods for model merging struggle to maintain performance gains as the number of merged models increases. |
| Approach: | They propose a Reparameterized Heavy-Tailed method to extend the merged model’s coverage and enhance performance. |
| Outcome: | The proposed method extends the merged model’s coverage and enhances performance on 19 benchmarks, including knowledge-intensive and general-purpose tasks. |
TOOL-ED: Enhancing Empathetic Response Generation with the Tool Calling Capability of LLM (2025.coling-main)
Copied to clipboard
| Challenge: | Empathetic conversation is a crucial characteristic in daily conversations between individuals. |
| Approach: | They propose an Emotional Knowledge Tool Calling framework which encapsulates commonsense knowledge bases as empathetic tools, enabling LLMs to integrate external knowledge flexibly. |
| Outcome: | The proposed framework can generate empathetic responses effectively on the TOOL-ED dataset. |
HiFT: A Hierarchical Full Parameter Fine-Tuning Strategy (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to fine-tuning language models use zeroth-order optimizers to conserve GPU memory. |
| Approach: | They propose a full-parameter fine-tuning strategy which updates a subset of parameters at each training step. |
| Outcome: | The proposed approach reduces the amount of gradients and optimizer state parameters residing in GPU memory at the same time, thereby reducing GPU memory usage. |
SAD: A Large-Scale Strategic Argumentative Dialogue Dataset (2026.acl-long)
Copied to clipboard
YongKang Liu, Jiayang Yu, Mingyang Wang, Yiqun Zhang, Ercong Nie, Shi Feng, Daling Wang, Kaisong Song, Hinrich Schuetze
| Challenge: | Argumentation is a key part of human reasoning and decision-making . existing argumentative corpora focus on single-turn settings, but multi-turn dialogues are often realized as multi-turned dialogues . |
| Approach: | They present a dataset for strategic multi-turn argumentation dialogues . they annotate each utterance with five strategy types, allowing multiple strategies per utterrance . |
| Outcome: | The proposed dataset shows that explicit prompting improves fluency, stylistic coherence and persuasiveness. |
LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing (2026.findings-acl)
Copied to clipboard
Hao Li, Yiqun Zhang, Zhaoyan Guo, Chenxu Wang, Shengji Tang, Qiaosheng Zhang, Yang Chen, Biqing Qi, Peng Ye, Lei Bai, Zhen Wang, Shuyue Hu
| Challenge: | Large language model (LLM) routing assigns each query to the best suitable model from an ensemble. |
| Approach: | They introduce a large-scale benchmark and unified framework for LLM routing . they find that many routing methods exhibit similar performance under unified evaluation . |
| Outcome: | The proposed benchmark provides comprehensive metrics for both performance-oriented and performance-cost trade-off routing. |
Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on hallucination detection for LLMs focus on how to identify possible factrelated errors in outputs. |
| Approach: | They propose an unsupervised training framework that leverages the internal states of LLMs for real-time hallucination detection without requiring manual annotations. |
| Outcome: | The proposed framework outperforms existing state-of-the-art methods in hallucination detection. |
LegalAgentBench: Evaluating LLM Agents in Legal Domain (2025.acl-long)
Copied to clipboard
Haitao Li, Junjie Chen, Jingli Yang, Qingyao Ai, Wei Jia, Youfeng Liu, Kai Lin, Yueyue Wu, Guozhi Yuan, Yiran Hu, Wuyue Wang, Yiqun Liu, Minlie Huang
| Challenge: | Existing general-domain benchmarks do not capture complexity of real-world judicial cognition and decision-making. |
| Approach: | They propose a benchmark specifically designed to evaluate LLM Agents in the legal domain. |
| Outcome: | The proposed benchmark includes 17 corpora from real-world legal scenarios and provides 37 tools for interacting with external knowledge. |
LoRaDA: Low-Rank Direct Attention Adaptation for Efficient LLM Fine-tuning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in parameter-efficient fine-tuning techniques allow for adjustments to only a minor fraction of the parameters of language models. |
| Approach: | They propose a low-rank direct attention adapted method for efficient LLM fine-tuning . they propose LMAM, which can bring negative attention to self-attention modules . |
| Outcome: | The proposed method outperforms the full fine-tuning method by 2.1% on GLUE benchmark. |
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities (2025.findings-acl)
Copied to clipboard
| Challenge: | Knowledge distillation (KD) with Direct Preference Optimization (DPO) has emerged as a promising approach to enhance the conversational abilities of smaller models using a larger teacher model. |
| Approach: | They propose a framework that integrates the teacher's distributional information into DPO distillation while preserving theoretical guarantees. |
| Outcome: | The proposed framework outperforms existing methods in restoring performance for pruned models and enhancing smaller models within the same LLM family. |
PRACTIQ: A Practical Conversational Text-to-SQL dataset with Ambiguous and Unanswerable Queries (2025.naacl-long)
Copied to clipboard
Mingwen Dong, Nischal Ashok Kumar, Yiqun Hu, Anuj Chauhan, Chung-Wei Hang, Shuaichen Chang, Lin Pan, Wuwei Lan, Henghui Zhu, Jiarong Jiang, Patrick Ng, Zhiguo Wang
| Challenge: | Existing text-to-SQL systems focus on user questions with clear intentions that can be answered, but real user questions can be ambiguous with multiple interpretations or unanswerable due to a lack of relevant data. |
| Approach: | They construct a conversational text-to-SQL dataset called PRACTIQ, consisting of ambiguous and unanswerable questions inspired by real-world user questions. |
| Outcome: | The proposed system generates conversations with four turns, generating the user’s question, an assistant response seeking clarification, and the user's clarified SQL response with the natural language explanation of the execution results. |
Cascaded Mutual Modulation for Visual Reasoning (D18-1)
Copied to clipboard
| Challenge: | Visual reasoning is a multi-step and compositional problem that requires intensive text-vision interactions. |
| Approach: | They propose a visual reasoning model that uses a feature-wise linear modulation technique to enable textual/visual pipelines to mutually control each other. |
| Outcome: | The proposed model outperforms existing models on visual reasoning benchmarks CLEVR and NLVR . it can generate a textual answer to a visual question answering problem with images . |
Design First, Code Later: Aesthetically Pleasing Template-Free Slides Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to producing presentation slides rely on fixed templates or executable code . Existing methods rely only on predefined templates and emit executable codes . |
| Approach: | They propose a hierarchical slides generation workflow DeepSlides that organizes slide design tasks without any predefined template or style. |
| Outcome: | The proposed framework outperforms baseline methods on evaluated metrics and achieves superior performance in human preference evaluations. |
Decoupling Reasoning and Knowledge Injection for In-Context Knowledge Editing (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing knowledge editing approaches directly edit model context without isolating target knowledge from the reasoning path of model inference, resulting in unreliable and low-quality outputs, especially in multi-hop tasks. |
| Approach: | They propose a framework that separates model reasoning from knowledge editing and propose 'DecKER' that allows users to modify specific factual associations without retraining the entire model. |
| Outcome: | The proposed framework significantly improves multi-hop reasoning performance by mitigating knowledge conflicts and preserving reasoning integrity. |
LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for cache compression are heuristic and lack dynamic budget allocation . cnn's john mccartney and johnny mccain present a new approach for cache eviction and dynamic budgets . |
| Approach: | They propose a unified framework for cache compression that minimizes information loss in transformer residual streams. |
| Outcome: | The proposed method consistently maintains top performance across task types. |
MTRouter: Cost-Aware Multi-Turn LLM Routing with History–Model Joint Embeddings (2026.acl-long)
Copied to clipboard
| Challenge: | Multi-turn, long-horizon tasks require dozens of sequential model calls per episode. |
| Approach: | They propose a cost-aware multi-turn LLM routing tool which encodes interaction history and candidate models into joint history–model embeddings and learns an outcome estimator from logged trajectories to predict turn-level model utility. |
| Outcome: | The proposed model reduces cost and performance by 58.7% on ScienceWorld and on Humanity’s Last Exam (HLE) and even reduces costs for held-out tasks. |
Knowledge Editing through Chain-of-Thought (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing knowledge editing methods focus on multi-hop QA tasks and require frequent retraining. |
| Approach: | They propose a new knowledge editing framework that updates large language models with new information to maintain their world knowledge without retraining. |
| Outcome: | The proposed method achieves state-of-the-art performance while offering superior generalization, effectiveness, and stability compared to existing methods. |