Papers by Zichen Zhu
META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI (2022.emnlp-main)
Copied to clipboard
| Challenge: | Current task-oriented dialogue systems focus on multi-turn text/speech interaction, then call back-end APIs to perform task. |
| Approach: | They propose a GUI-based task-oriented dialogue system that can perform GUI operations on real APPs without invoking TOD-specific backend APIs. |
| Outcome: | The proposed GUI-based task-oriented dialogue system can perform GUI operations on real APPs and execute tasks without invoking TOD-specific backend APIs. |
Alignment for Efficient Tool Calling of Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in tool learning have enabled large language models to integrate external tools, enhancing their task performance by expanding their knowledge boundaries. |
| Approach: | They propose a framework that combines probabilistic knowledge boundary estimation with dynamic decision-making to allow LLMs to better assess when to invoke tools based on their confidence. |
| Outcome: | The proposed framework shows significant improvements in tool efficiency by reducing unnecessary tool usage. |
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding (2026.findings-acl)
Copied to clipboard
Junxi Wang, Te Sun, Jiayi Zhu, Junxian Li, Haowen Xu, Zichen Wen, Xuming Hu, Zhiyu li, Linfeng Zhang
| Challenge: | StreamMeCo is an efficient Stream Agent Memory Compression framework for video understanding. |
| Approach: | They propose an efficient Stream Agent Memory Compression framework that evicts redundant memory nodes and introduces a time-decay memory retrieval mechanism to mitigate performance degradation. |
| Outcome: | The proposed framework achieves 1.87 speedup in memory retrieval while delivering an average accuracy improvement of 1.0% on three challenging benchmark datasets. |
Black-Box Membership Inference Attacks for Video Training Data in Multimodal Large Language Models (2026.acl-long)
Copied to clipboard
Jinrui Wang, Zhenfeng Gao, Wendan Wang, Huili Wang, Zichen Qin, Linjie Zhu, Hongke Fu, Shangguang Wang, Tao Qi
| Challenge: | Existing methods assess model memorization of key semantic concepts within a video but do not provide reliable evidence that a specific video was used during training. |
| Approach: | They propose a black-box MIA framework that can provide reliable evidence of specific video data usage for training multimodal large language models. |
| Outcome: | The proposed framework can provide reliable evidence of specific video data usage for training multimodal large language models. |
ChatKBQA: A Generate-then-Retrieve Framework for Knowledge Base Question Answering with Fine-tuned Large Language Models (2024.findings-acl)
Copied to clipboard
Haoran Luo, Haihong E, Zichen Tang, Shiyao Peng, Yikai Guo, Wentai Zhang, Chenghao Ma, Guanting Dong, Meina Song, Wei Lin, Yifan Zhu, Anh Tuan Luu
| Challenge: | Existing KBQA methods address inefficient knowledge retrieval and semantic parsing errors. |
| Approach: | They propose a generatethen-retrieve KBQA framework that generates logical form and replaces entities and relations with an unsupervised retrieval method to improve both generation and retrieval more directly. |
| Outcome: | Experimental results show that ChatKBQA achieves new state-of-the-art performance on standard KBQA datasets, WebQSP, and CWQ. |