Papers by Xiaoxin Lu
SmartBench: Is Your LLM Truly a Good Chinese Smartphone Assistant? (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for Large Language Models focus on objective tasks like mathematics and coding in English, which do not reflect the practical use cases of on-device LLMs in real-world mobile scenarios. |
| Approach: | They propose a benchmark to evaluate the capabilities of on-device Large Language Models in Chinese mobile contexts. |
| Outcome: | The proposed framework evaluates on-device LLMs and MLLMs in Chinese . it provides a standardized framework for evaluating LLM performance on real smartphones . |
Fair Abstractive Summarization of Diverse Perspectives (2024.naacl-long)
Copied to clipboard
Yusen Zhang, Nan Zhang, Yixin Liu, Alexander Fabbri, Junru Liu, Ryo Kamoi, Xiaoxin Lu, Caiming Xiong, Jieyu Zhao, Dragomir Radev, Kathleen McKeown, Rui Zhang
| Challenge: | Existing work on summarization metrics and large language models has not explored fair abstractive summarizing. |
| Approach: | They propose four reference-free automatic metrics to measure the differences between target and source perspectives. |
| Outcome: | The proposed methods alleviate fair abstractive summarization on user-generated data. |
Doctor Recommendation in Online Health Forums via Expertise Learning (2022.acl-long)
Copied to clipboard
| Challenge: | Currently, manual doctor allocations are used to handle large volumes of queries, limiting the efficiency to help patients in sheer quantities. |
| Approach: | They propose to use patient queries to model doctor recommendation using their profiles and past dialogues to estimate their capabilities. |
| Outcome: | The proposed model outperforms baseline models on a Chinese online health forum, outperforming baseline models. |
Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on textual plan generation only focus on LLMs, enabling applications in robotics, virtual assistants, and instruc. |
| Approach: | They propose a framework that generates and refines text-image plans step-by-step . they collect a new benchmark consisting of 1,100 tasks and their text- image pair solutions covering 11 daily topics. |
| Outcome: | The proposed framework generates and refines text-image plans step-by-step and improves on existing models. |