Papers by Zhenran Xu
ExplainCPE: A Free-text Explanation Benchmark of Chinese Pharmacist Examination (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing explanation datasets for large language models are limited to the English language and general domain, leading to a scarcity of linguistic diversity and a lack of resources in specialized domains, such as medical. |
| Approach: | They propose to use a medical dataset to assess the interpretability of Large Language Models (LLMs) . they propose to analyze medical text and generate rationales for their decisions . |
| Outcome: | The proposed model passes the pharmacist examination with a 75.7% accuracy, while other models like ChatGPT fail. |
ComfyUI-R1: Exploring Reasoning Models for Workflow Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | ComfyUI-R1 is the first large reasoning model for automated workflow generation. |
| Approach: | They propose a large reasoning model for automated workflow generation that builds on curated knowledge bases and a two-stage framework to fine-tune models for cold start and reinforcement learning for incentivizing reasoning capability. |
| Outcome: | The proposed model achieves 97% format validity rate, high pass rate, node-level and graph-level F1 scores, surpassing prior state-of-the-art methods that employ leading closed-source models such as GPT-4o and Claude series. |
Generative Multimodal Entity Linking (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing Entity Linking methods focus on designing complex multimodal interaction mechanisms and require fine-tuning all model parameters. |
| Approach: | They propose a framework for multimodal entity linking based on Large Language Models (LLMs) that trains a feature mapper to enable cross-modal interactions. |
| Outcome: | The proposed framework achieves state-of-the-art on two well-established datasets with a performance gain of 7.7% on WikiDiverse and 8.8% on Wikileaks. |
MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation methods for complex multi-shot video are anchored to single-shot paradigms, lacking comprehensive story assets and cross-shot metrics. |
| Approach: | They propose a framework that synergizes the high-level semantic reasoning of Large Multimodal Models with the fine-grained perceptual rigor of domain-specific expert models. |
| Outcome: | The proposed framework synergizes the high-level semantic reasoning of Large Multimodal Models with the fine-grained perceptual rigor of domain-specific expert models. |
ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development (2025.acl-demo)
Copied to clipboard
Zhenran Xu, Yangxue Yangxue, Yiyu Wang, Qingli Hu, Zijiao Wu, Baotian Hu, Longyue Wang, Weihua Luo, Kaifu Zhang
| Challenge: | ComfyUI-Copilot is a large language model-powered plugin for AI-driven art creation. |
| Approach: | They propose a large language model-powered plugin to enhance the usability of ComfyUI. |
| Outcome: | The new plugin improves the usability and efficiency of ComfyUI . it offers intelligent node and model recommendations and automated one-click workflow construction. |
ComfyFlow: Benchmarking LLMs for AIGC Workflow Generation (2026.findings-acl)
Copied to clipboard
Zhenran Xu, Yiyu Wang, Yunxin li, Muyang Ye, null Yangxue, Kai Chen, Longyue Wang, Weihua Luo, Baotian Hu, Min Zhang
| Challenge: | Large language models (LLMs) have shown promising advances in tackling human-level tasks, but generating workflows for collaborative AI systems remains a critical and challenging step. |
| Approach: | They propose a benchmark to evaluate LLMs’ ability to generate executable and instruction-following AIGC workflows in ComfyUI. |
| Outcome: | The proposed benchmarks show that LLMs can generate executable and instruction-following AIGC workflows in ComfyUI. |
Revisiting Sparse Retrieval for Few-shot Entity Linking (2023.emnlp-main)
Copied to clipboard
| Challenge: | Entity linking (EL) aims to link ambiguous mentions to their corresponding entities in a knowledge base. |
| Approach: | They propose an ELECTRA-based keyword extractor to denoise the mention context and construct a better query expression. |
| Outcome: | The proposed method outperforms state-of-the-art models on the ZESHEL dataset by a significant margin. |
WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing GUI agents perform poorly on multi-application tasks, stalling at early sub-goals. |
| Approach: | They propose to assess GUI Agents on complex multi-step tasks that mirror real-world professions. |
| Outcome: | The proposed benchmark contains 181 tasks with an average of 5.0 sub-goals across 17 common desktop applications, of which 78% are inherently multi-application. |
A Read-and-Select Framework for Zero-shot Entity Linking (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods focus on the candidate retrieval stage and ignore the essential candidate ranking stage, which disambiguates among entities and makes the final linking prediction. |
| Approach: | They propose a read-and-select framework that models the main components of entity disambiguation . they use mention context to output mention-aware entity representations . |
| Outcome: | The proposed framework achieves state-of-the-art performance on established zero-shot entity linking dataset ZESHEL with 2.55% micro-average accuracy gain, with no need for laborious multi-phase pre-training used in most of the previous work. |
MultiSkill: Evaluating Large Multimodal Models for Fine-grained Alignment Skills (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evaluation settings for large multimodal models focus on coarse-grained evaluation without considering skill composition required by specific instructions. |
| Approach: | They propose an evaluation protocol that assesses large multimodal models across multiple fine-grained skills for alignment with human values. |
| Outcome: | The proposed evaluation protocol decomposes coarse-level scoring to fine-grained skill set-level score tailored to each instruction. |
A Unified Agentic Framework for Evaluating Conditional Image Generation (2025.acl-long)
Copied to clipboard
Jifang Wang, Yangxue Yangxue, Longyue Wang, Zhenran Xu, Yiyu Wang, Yaowei Wang, Weihua Luo, Kaifu Zhang, Baotian Hu, Min Zhang
| Challenge: | Conditional image generation is a popular and personalization-oriented task, but there are challenges in developing task-agnostic, reliable, and explainable evaluation metrics. |
| Approach: | They propose a unified agentic framework for comprehensive evaluation of conditional image generation tasks. |
| Outcome: | The proposed framework achieves a high correlation with human assessments on seven prominent image generation tasks. |