Papers by Jingtao Cao
AppBench: Planning of Multiple APIs from Various APPs for Complex User Instruction (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing state-of-the-art Large Language Models (LLMs) still cannot perform well in this situation even with the help of in-context learning and finetuning. |
| Approach: | They propose a benchmark to evaluate LLMs’ ability to plan and execute multiple APIs from various sources in order to complete the user’s task. |
| Outcome: | The proposed benchmarks show that the existing state-of-the-art LLMs still cannot perform well in this situation even with in-context learning and finetuning. |
GUI0: Self-Evolving Foundational GUI Agents in Super App Ecosystems (2026.acl-long)
Copied to clipboard
Xinyi Wang, Wei Dai, Kyle Qiao, Ke Wang, Peng Chen, Gang Cao, null Kangqin, Zhongpu Wang, Xiaode Zhang, Yanming Liu, Jihao Gu, Jingtao Xu, Gong Zhi
| Challenge: | Automated interaction with graphical user interfaces (GUIs) is central to general artificial intelligence, but remains challenging within Super App ecosystems. |
| Approach: | They propose a framework synergizing autonomous data synthesis with dual-agent co-evolution . GUI0 establishes a domain-aware foundation model via synthesized corpora and employs curriculum-driven reinforcement learning . |
| Outcome: | The proposed framework outperforms Gemini-2.5-Pro and Claude-4-Sonnet in the SuperAPP benchmark and has universal efficacy across base models. |
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing metrics, such as CLIP, measure the semantic alignment between single prompts and their corresponding images, but they fail to evaluate a model’s generalizability across a broad spectrum of textual inputs. |
| Approach: | They propose a metric that leverages the power of Large Language Models to sample from the visual text domain and assess its generalizability. |
| Outcome: | The proposed metric evaluates the generalizability of T2I models and provides valuable insights during the finetuning process. |