Papers by Zeming Liu
RETAIL: Towards Real-world Travel Planning for Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing travel planning systems assume users provide explicit queries, limiting their practical utility. |
| Approach: | They propose a dataset RETAIL which supports decision-making for implicit queries while covering explicit queries. |
| Outcome: | The proposed model achieves a 1.0% pass rate, suggesting real-world travel planning remains challenging. |
AppBench: Planning of Multiple APIs from Various APPs for Complex User Instruction (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing state-of-the-art Large Language Models (LLMs) still cannot perform well in this situation even with the help of in-context learning and finetuning. |
| Approach: | They propose a benchmark to evaluate LLMs’ ability to plan and execute multiple APIs from various sources in order to complete the user’s task. |
| Outcome: | The proposed benchmarks show that the existing state-of-the-art LLMs still cannot perform well in this situation even with in-context learning and finetuning. |
RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models (2025.findings-emnlp)
Copied to clipboard
Jingjing Liu, Zeming Liu, Zihao Cheng, Mengliang He, Xiaoming Shi, Yuhang Guo, Xiangrong Zhu, Yuanfang Guo, Yunhong Wang, Haifeng Wang
| Challenge: | Large Language Models (LLMs) have exhibited significant proficiency in code debugging, especially in automatic program repair. |
| Approach: | They propose a repository-level code debugging dataset with 22 subtypes of errors that supports 8 commonly used programming languages and 3 debug tasks. |
| Outcome: | The proposed dataset supports 8 commonly used programming languages and 3 debugging tasks. |
FAME: Towards Factual Multi-Task Model Editing (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models embed extensive knowledge and perform exceptionally well across tasks. outdated knowledge or factual errors within LLMs can lead to misleading or incorrect responses. |
| Approach: | They propose to use a dataset to enhance the practicality of model editing to correct inaccurate information within LLMs. |
| Outcome: | The proposed method performs excellently across tasks and scenarios, confirming its practicality. |
Weak2Wise: An Automated, Lightweight Framework for Weak-LLM-Friendly Reasoning Synthesis (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to finetuning large language models rely on expensive manual annotations or auxiliary models and fail to address the unique constraints of smaller "weak" LLMs. |
| Approach: | Weak2Wise is a fully automated framework for synthesizing highquality, weak-LLM-friendly reasoning traces. |
| Outcome: | Weak2Wise is a fully automated, lightweight framework for synthesizing highquality, weak-LLM-friendly reasoning traces. |
TED-EL: A Corpus for Speech Entity Linking (2024.lrec-main)
Copied to clipboard
| Challenge: | Current entity linking tasks rely on textual information, but entities usually exist in textual, audio, and visual contexts in real-world data such as social media and video websites. |
| Approach: | They propose a speech entity linking task to recognize mentions from speech and link them to entities in knowledge bases. |
| Outcome: | The proposed model outperforms the existing models on the TED-EL dataset, scoring an F1 score of 60.68%. |
TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments (2025.findings-acl)
Copied to clipboard
Yuheng Lu, Qian Yu, Hongru Wang, Zeming Liu, Wei Su, Yanping Liu, Yuhang Guo, Maocheng Liang, Yunhong Wang, Haifeng Wang
| Challenge: | Existing GUI agents struggle to adapt to dynamic and interconnected nature of real-world digital environments, authors show . |
| Approach: | They propose a benchmark to evaluate the transferability of GUI agents across three key dimensions . transBench includes 15 app categories with diverse functionalities . |
| Outcome: | The proposed benchmark shows that existing GUI agents struggle to adapt to dynamic, interconnected environments. |
ToolSpectrum: Towards Personalized Tool Utilization for Large Language Models (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches focus on functional tool selection following user instructions while overlooking the critical role of context-aware personalization in tool selection. |
| Approach: | They propose a benchmark to evaluate LLMs’ capabilities in personalized tool utilization. |
| Outcome: | The proposed benchmark evaluates LLMs' capabilities in personalized tool utilization. |
GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art (2025.acl-long)
Copied to clipboard
Yiming Lei, Chenkai Zhang, Zeming Liu, Haitao Leng, ShaoGuo Liu, Tingting Gao, Qingjie Liu, Yunhong Wang
| Challenge: | Existing benchmarks for video comment art are constrained by their limited modalities and insufficient categories, hindering creativity in video-based comment art creation. |
| Approach: | They propose a benchmark that integrates video and text modalities to evaluate MLLMs’ abilities to compose video Comment art. |
| Outcome: | The proposed framework integrates video and text modalities to evaluate MLLMs’ abilities to compose video comment art. |
Automatic Evaluate Dialogue Appropriateness by Using Dialogue Act (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluations of dialogue quality rely on human judgments, which are time-consuming, labor-intensive, prone to biases, and lacking objectivity. |
| Approach: | They propose a method that utilizes the underlying patterns of dialogue act transitions to evaluate the appropriateness of chatbot responses. |
| Outcome: | The proposed method proves that human judgments are time-consuming, labor-intensive, and lacking objectivity. |
Exploring In-Image Machine Translation with Real-World Background (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing models for IIMT focus on simplified scenarios, which is far from reality and impractical for applications in the real world. |
| Approach: | They propose a model that separates the background and text-image from the source image and performs translation on the text- image directly. |
| Outcome: | The proposed model improves translation quality and visual effect in complex scenarios . it separates background and text-image from source image and performs translation on the text- image directly . |
PRIM: Towards Practical In-Image Multilingual Machine Translation (2025.emnlp-main)
Copied to clipboard
| Challenge: | Current research on in-image machine translation focuses on synthetic data with simple background, single font, fixed text position, and bilingual translation. |
| Approach: | They propose an end-to-end model to handle the challenge of practical conditions in PRIM . they annotate a real-world one-line text image with complex background, fonts, diverse text positions . |
| Outcome: | The proposed model improves translation quality and visual effect compared to other models. |
DocMEdit: Towards Document-Level Model Editing (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing models only output short phrases or sentences, raising doubts about their practical usability. |
| Approach: | They propose a dataset focused on document-level model editing that aims to correct errors and outdated knowledge in Large language models (LLMs) they propose to use document-based model editing to improve model capabilities in real-world scenarios. |
| Outcome: | The proposed model editing task improves model capabilities in real-world scenarios and reduces the cost of retraining. |
HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices (2025.acl-long)
Copied to clipboard
| Challenge: | Existing state-of-the-art LLMs cannot perform well in situations where instructions are invalid or multiple devices are involved. |
| Approach: | They propose to integrate large language models into smart home assistants by enhancing their ability to accurately understand user needs and respond appropriately. |
| Outcome: | The proposed dataset is the first with valid and invalid instructions across devices . it achieves only 0.0% success rate in the scenario of invalid multi-device instructions . |
KwaiChat: A Large-Scale Video-Driven Multilingual Mixed-Type Dialogue Corpus (2025.findings-naacl)
Copied to clipboard
Xiaoming Shi, Zeming Liu, Yiming Lei, Chenkai Zhang, Haitao Leng, Chuan Wang, Qingjie Liu, Wanxiang Che, Yunhong Wang
| Challenge: | Currently, video-based dialogue systems rely on a single dialogue type, hindering their versatility in practical applications. |
| Approach: | They propose to generate video-driven multilingual mixed-type dialogues using KwaiChat . they propose to create a video-based multilingual mix of 4 dialogue types, 30 domains, 4 languages, 13 topics . |
| Outcome: | The proposed model performs best on KwaiChat but is not perfect in this situation. |
Stealthy Jailbreak Attacks on Large Language Models via Benign Data Mirroring (2025.naacl-long)
Copied to clipboard
Honglin Mu, Han He, Yuxin Zhou, Yunlong Feng, Yang Xu, Libo Qin, Xiaoming Shi, Zeming Liu, Xudong Han, Qi Shi, Qingfu Zhu, Wanxiang Che
| Challenge: | Existing black-box jailbreak methods often rely on model feedback . existing methods may be intercepted by content moderators during the search process . |
| Approach: | They propose a method that guides malicious prompt construction by local training a mirror model of the target black-box model through benign data distillation. |
| Outcome: | The proposed method achieves a 92% attack success rate and 80% stealth rate on a subset of AdvBench. |
Where to Go for the Holidays: Towards Mixed-Type Dialogs for Clarification of User Goals (2022.acl-long)
Copied to clipboard
| Challenge: | a dialog system posits that users have figured out clear and specific goals . but in many real-world scenarios, users struggle to figure out specific goals by determining all the necessary slots. |
| Approach: | They propose a mixed-type dialog model with a Prompt-based continual learning mechanism . they collect 5k dialog sessions and 168k utterances for 4 dialog types and 5 domains . |
| Outcome: | The proposed model provides user-goal-related knowledge to help figure out clear and specific goals . it can be extended to any specific type by utilizing existing dialog corpora effectively. |
Beyond Unimodal Shortcuts: MLLMs as Cross-Modal Reasoners for Grounded Named Entity Recognition (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to GMNER use MLLMs as auxiliary tools, causing cumulative error propagation and a lack of rigorous cross-modal verification. |
| Approach: | They propose a model that enforces structured cross-modal reasoning through Multi-style Reasoning Schema Injection and Constraint-guided Verifiable Optimization. |
| Outcome: | The proposed model enforces structured cross-modal reasoning through multi-style Reasoning Schema Injection and Constraint-guided Verifiable Optimization. |
PEC-Home: Interpretation of Progressively Elliptical Commands in Smart Homes (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing home assistants struggle to interpret elliptical commands based on ellipine expressions . current assistants overlook the progressive omission that occurs in human dialogue as context accumulates - limiting their effectiveness in real-world applications . |
| Approach: | They propose a simulated home dataset specifically designed for interpreting progressively elliptical commands in smart homes. |
| Outcome: | The proposed dataset shows that existing home assistants struggle to execute user-intended operations based solely on elliptical commands. |
Self-Reasoning Language Models: Unfold Hidden Reasoning Chains with Few Reasoning Catalyst (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent studies have demonstrated that inference-time scaling increases performance of Large Language Models (LLMs) in various reasoning tasks such as mathematics and complex question answering by increasing the length of Chain-of-Thought (CoT). |
| Approach: | They propose a model which synthesizes longer CoT data and iteratively improves performance through self-training by incorporating a few demonstration examples. |
| Outcome: | The proposed model achieves an average improvement of more than +2.5 points across five reasoning tasks: MMLU, GSM8K, ARC-C, HellaSwag, and BBH on two backbone models. |
Live-Aid: A Large-Scale Dialogue Dataset and Benchmark for Interleaved Multi-party Interactions in Live Streaming (2026.findings-acl)
Copied to clipboard
Yiming Lei, Yize Fan, Zeming Liu, Jiaji Dong, Hui Qiu, Haitao Leng, Qingjie Liu, Kehai Chen, Tingting Gao, Yunhong Wang
| Challenge: | Existing Multimodal Large Language Models struggle with dynamic interactions due to the scarcity of high-quality interleaved data. |
| Approach: | They propose a large-scale interleaved live interaction Chinese dataset with human-annotated video responses. |
| Outcome: | The proposed model can be used to evaluate live interactions in Chinese over 1,100 hours and 80,037 dialogue turns. |
In-Image Neural Machine Translation with Segmented Pixel Sequence-to-Sequence Model (2023.findings-emnlp)
Copied to clipboard
| Challenge: | In-Image Machine Translation (IIMT) aims to convert images containing texts from one language to another. |
| Approach: | They propose an end-to-end model instead of the traditional cascade methods which use optical character recognition followed by neural machine translation and text rendering. |
| Outcome: | The proposed model outperforms both cascade methods and current model in translation quality and robustness across various dimensions. |
Towards Conversational Recommendation over Multi-Type Dialogs (2020.acl-main)
Copied to clipboard
| Challenge: | In recent years, there has been a significant increase in the work of conversational recommendation due to the rise of voice-based bots. |
| Approach: | They use a Chinese dialog dataset DuRecDial to study conversational recommendation in the context of multi-type dialogs where bots can proactively lead a conversation from a non-recommendation dialog to a recommendation dialog. |
| Outcome: | The proposed dataset allows to investigate different parts of the overall problem, e.g., how to naturally lead a dialog, how interact with users for recommendation. |
Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have advanced machine translation (MT) a meta-evaluation dataset focused on non-literal translations is lacking . experimental results show the inaccuracies of traditional MT metrics and the limitations of LLM-as-a-Judge. |
| Approach: | They propose a meta-evaluation framework that leverages sub-agents to evaluate machine translation metrics. |
| Outcome: | The proposed framework improves on the knowledge cutoff and score inconsistency problem. |
Mem2Evolve: Towards Self-Evolving Agents via Co-Evolutionary Capability Expansion and Experience Distillation (2026.acl-long)
Copied to clipboard
Zihao Cheng, Zeming Liu, Yingyu Shan, Xinyi Wang, Xiangrong Zhu, Yunpu Ma, Hongru Wang, Yuhang Guo, Wei Lin, Yunhong Wang
| Challenge: | Existing frameworks that focus on static tools and static assets are ineffective for self-evolving agents. |
| Approach: | They propose a paradigm of co-evolutionary Capability Expansion and Experience Distillation that leverages accumulated experience to guide dynamic creation of assets. |
| Outcome: | The proposed framework improves performance in single-task and cross-task settings by 18.53% over standard LLMs, 11.80% over agents evolving solely through experience, and 6.46% over those evolving solelly through asset creation. |
PEAP: Proactive Embodied Action Sequence Planning with Joint Understanding of Vision and Audio Perception (2026.acl-long)
Copied to clipboard
| Challenge: | Embodied action sequence planning focuses on the capability of embodied agents to implement action planning via environmental perception without explicit human instructions. |
| Approach: | They propose to use a multimodal dataset to evaluate the performance of multiple large language models to evaluate their models' environmental perception capabilities. |
| Outcome: | The proposed model shows that it lacks accurate environmental perception capabilities and that it can improve on the PEAP dataset. |
DuRecDial 2.0: A Bilingual Parallel Corpus for Conversational Recommendation (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets for conversational recommendation are limited to English and Chinese . |
| Approach: | They propose a bilingual parallel human-to-human recommendation dialog dataset . the data item is annotated in two languages, both English and Chinese . |
| Outcome: | The proposed dataset provides a testbed for future studies of multilingual and cross-lingual conversational recommendation. |
Medical Dialogue System: A Survey of Categories, Methods, Evaluation and Challenges (2024.findings-acl)
Copied to clipboard
Xiaoming Shi, Zeming Liu, Li Du, Yuxuan Wang, Hongru Wang, Yuhang Guo, Tong Ruan, Jie Xu, Xiaofan Zhang, Shaoting Zhang
| Challenge: | Existing medical dialogue systems have significant potential to simplify diagnostic procedure and reduce the cost of collecting information from patients. |
| Approach: | They analyze 325 papers from well-known computer science, natural language processing conferences and journals to find out the major challenges of medical dialog systems. |
| Outcome: | The proposed systems have been surveyed in the medical community but have not been evaluated from a technical perspective. |
SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches fail to fully capture all risks in tool utilization, resulting in financial loss or privacy leaking. |
| Approach: | They propose a framework to assess the safety of LLM tool utilization in a prospective manner, covering malicious user instructions and diverse practical toolsets. |
| Outcome: | The proposed framework significantly enhances LLMs’ self-awareness, enabling a more safer and trustworthy tool utilization. |
MidMed: Towards Mixed-Type Dialogues for Medical Consultation (2023.acl-long)
Copied to clipboard
| Challenge: | Current medical dialogue systems assume that patients have explicit goals but are often unavailable in real-world situations due to the lack of medical knowledge. |
| Approach: | They propose a human-to-human mixed-type medical consultation dialogue corpus . they build benchmarking baselines on MidMed and propose an instruction-guiding framework . Experimental results show the effectiveness of InsMed . |
| Outcome: | The proposed system can help patients clarify their goals in real-world situations . it covers four departments with 8,309 dialogues and provides benchmarking baselines . |
Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation (2025.findings-acl)
Copied to clipboard
Qiyue Gao, Xinyu Pi, Kevin Liu, Junrong Chen, Ruolan Yang, Xinqi Huang, Xinyu Fang, Lu Sun, Gautham Kishore, Bo Ai, Stone Tao, Mengyang Liu, Jiaxi Yang, Chao-Jung Lai, Chuanyang Jin, Jiannan Xiang, Benhao Huang, Zeming Chen, David Danks, Hao Su, Tianmin Shu, Ziqiao Ma, Lianhui Qin, Zhiting Hu
| Challenge: | Recent studies have evaluated and shown limitations in specific capabilities such as visual understanding, but a systematic evaluation of VLMs’ fundamental WM abilities remains absent. |
| Approach: | They propose a framework that assesses perception and prediction to provide an atomic evaluation of VLMs as WMs. |
| Outcome: | The proposed framework assesses perception and prediction abilities on 15 latest VLMs and compares them to human-level models. |
XDailyDialog: A Multilingual Parallel Dialogue Corpus (2023.acl-long)
Copied to clipboard
Zeming Liu, Ping Nie, Jie Cai, Haifeng Wang, Zheng-Yu Niu, Peng Zhang, Mrinmaya Sachan, Kaiping Peng
| Challenge: | Existing datasets for open-domain dialogue modeling limited to a single language . absence of multilingual datasets hinders development of robust open- domain dialog systems . |
| Approach: | They propose a multilingual parallel open-domain dialog dataset to explore multilingual and cross-lingual open- domain dialog. |
| Outcome: | The proposed model can be used to explore multilingual and cross-lingual open-domain dialogs in other languages. |
Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling (2025.acl-long)
Copied to clipboard
| Challenge: | Current error-handling works are performed in a passive manner, with explicit error- handling instructions. |
| Approach: | They propose a new benchmark to analyze LLMs' performance on a mis-prompt benchmark and a dataset to promote further research. |
| Outcome: | The proposed benchmark shows that current LLMs show poor performance on proactive error handling, and that SFT improves on error handling instances. |
Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges (2025.findings-acl)
Copied to clipboard
Hongru Wang, Wenyu Huang, Yufei Wang, Yuanhao Xi, Jianqiao Lu, Huan Zhang, Nan Hu, Zeming Liu, Jeff Z. Pan, Kam-Fai Wong
| Challenge: | Existing benchmarks that assess Language Models (LMs) as Language Agents (LAs) for tool use focus on stateless, single-turn interactions or partial evaluations, overlooking the inherent stateful nature of interactions in multi-turn applications. |
| Approach: | They propose a multi-turn dialogue dataset with stateful tool interactions considering the whole life cycle of tool use across six key tasks in three stages . they also build VirtualMobile – an embodied virtual mobile evaluation environment to simulate API calls and assess the robustness of the created APIs. |
| Outcome: | The proposed dataset evaluates 13 open- and closed-source LLMs and provides detailed analysis at each stage. |
False Sense of Security: Why Probing-based Malicious Input Detection Fails to Generalize (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent work has leveraged probing-based approaches to study the separability of malicious and benign inputs in Large Language Models’ internal representations. |
| Approach: | They propose to use probing-based methods to study separability of malicious and benign inputs in LLMs' internal representations to detect harmful and benign content. |
| Outcome: | The proposed methods show that they learn superficial patterns rather than semantic harmfulness. |
Deterministic Reversible Data Augmentation for Neural Machine Translation (2024.findings-acl)
Copied to clipboard
| Challenge: | Recent neural machine translation models have improved translation quality but they also introduce small perturbations like misspelling and paraphrasing. |
| Approach: | They propose a method that generates multi-granularity subword representations with reversible operations and deterministic segmentations. |
| Outcome: | The proposed method outperforms strong baselines on several translation tasks with a clear margin and exhibits good robustness in noisy, low-resource, and cross-domain datasets. |
Flow2Code: Evaluating Large Language Models for Flowchart-based Code Generation Capability (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing code generation benchmarks neglect flowchart-based code generation . existing benchmarks lack flowcharting-based evaluation, limiting the potential of large language models and minimizing human error. |
| Approach: | They propose to use flowcharts to evaluate existing LLMs' code generation capabilities. |
| Outcome: | The proposed benchmarks show that the supervised fine-tuning technique contributes greatly to the models’ performance. |