InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection (2026.eacl-long)
Copied to clipboard
Yuhang Liu, Pengxiang Li, Zishu Wei, Congkai Xie, Xueyu Hu, Xinchen Xu, Shengyu Zhang, Xiaotian Han, Hongxia Yang, Fei Wu
| Challenge: | Existing GUI Agents face challenges in multi-step reasoning and reliance on textual annotations, limiting their effectiveness. |
| Approach: | They propose an MLLM-based GUI Agent with a two-stage supervised fine-tuning pipeline that enhances GUI understanding and grounding. |
| Outcome: | InfiGUIAgent achieves competitive performance on several GUI benchmarks, highlighting the impact of native reasoning skills in enhancing GUI interaction for automation tasks. |
Similar Papers
INREACT: An Inspire-Then-Reinforce Training Framework For Multimodal GUI Agent (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing multimodal large language models struggle with precise localization of small elements. |
| Approach: | They propose a multimodal GUI agent framework that unifies observing, thinking, and acting for precise and interpretable decision-making. |
| Outcome: | The proposed framework unifies observing, thinking, and acting for precise and interpretable decision-making. |
CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation (2024.findings-acl)
Copied to clipboard
| Challenge: | Current vital challenges for autonomous agents lie in two aspects: dependence on strong (M)LLMs and insufficient GUI environment modeling. |
| Approach: | They propose a comprehensive cognitive LLM agent with two novel approaches to improve GUI automation performance. |
| Outcome: | The proposed agent achieves state-of-the-art performance on AITW and META-GUI benchmarks. |
You Only Look at Screens: Multimodal Chain-of-Action Agents (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to creating autonomous graphical user interfaces rely on external tools and application-specific APIs to interpret the environment. |
| Approach: | They propose a multimodal solution that directly interacts with the user interface without environment parsing. |
| Outcome: | The proposed solution bypasses environment parsing and reliance on application-dependent APIs. |
Towards Scalable Lightweight GUI Agents via Multi-role Orchestration (2026.findings-acl)
Copied to clipboard
Ziwei Wang, Junjie Zheng, Leyang Yang, Sheng Zhou, Xiaoxuan Tang, Fang Zhouhua, Zhiwei Liu, Dajun Chen, Yong Li, Jiajun Bu
| Challenge: | Advanced GUI agents suffer from prohibitive deployment costs on resource-constrained devices. |
| Approach: | They propose a lightweight GUI agent with GUI-specific knowledge and task scalability . LAMO-3B supports monolithic execution and MAS-style orchestration . |
| Outcome: | The proposed GUI agent LAMO-3B supports monolithic execution and MAS-style orchestration. |
AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning (2025.emnlp-demos)
Copied to clipboard
Zhong Zhang, Yaxi Lu, Yikun Fu, Yupeng Huo, Shenzhi Yang, Yesai Wu, Han Si, Xin Cong, Haotian Chen, Yankai Lin, Xie Xie, Wei Zhou, Wang Xu, Zhou Su, Zhongwu Zhai, Xiaoming Liu, null Meiyudong, Jianming Xu, Hongyan Tian, Chongyi Wang, Chi Chen, Yuan Yao, Zhiyuan Liu, Maosong Sun
| Challenge: | Large language model agents have enabled GUI-based automation, but their deployment is limited by noisy data, poor generalization, and lack of support for non-English GUIs. |
| Approach: | They propose an 8B-parameter GUI agent built for robust and efficient on-device GUI interaction. |
| Outcome: | The proposed GUI agent achieves promising performance on five public benchmarks and proposed Chinese benchmark CAGUI. |
AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations (2025.acl-long)
Copied to clipboard
| Challenge: | State-of-the-art multimodal web agents can perform many web tasks by processing user instructions and interacting with graphical user interfaces (GUIs). |
| Approach: | They propose to build multimodal web agents for few-shot adaptability using human demonstrations to improve their generalization and adaptability. |
| Outcome: | The proposed framework enables both proprietary and open-weights multimodal web agents to adapt to new websites and domains using few human demonstrations. |
Retrieval-augmented GUI Agents with Generative Guidelines (2025.emnlp-main)
Copied to clipboard
| Challenge: | GUI agents powered by vision-language models struggle with real-world tasks due to their complex nature and limited training data. |
| Approach: | They propose a lightweight vision-language model that leverages web tutorials at inferencetime to synthesize GUI agents. |
| Outcome: | The proposed agent outperforms baseline GUI agents and surpasses other inference baselines by 2.6% to 13.3% across two model sizes. |
GUICourse: From General Vision Language Model to Versatile GUI Agent (2025.acl-long)
Copied to clipboard
Wentong Chen, Junbo Cui, Jinyi Hu, Yujia Qin, Junjie Fang, Yue Zhao, Chongyi Wang, Jun Liu, Guirong Chen, Yupeng Huo, Yuan Yao, Yankai Lin, Zhiyuan Liu, Maosong Sun
| Challenge: | Graphical User Interfaces (GUIs) are a pivotal medium for human-computer interaction. |
| Approach: | They propose a series of datasets for training visual-based GUI agents using general VLMs. |
| Outcome: | The proposed GUICourse datasets show that even a small-sized GUI agent performs better on GUI tasks. |
Ponder & Press: Advancing Visual GUI Agent towards General Computer Control (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing multimodal large language models (MLLMs) lack visual inputs to ground objects, limiting flexibility across diverse software environments and platforms. |
| Approach: | They propose a divide-and-conquer framework for general computer control that uses only visual inputs to create a purely human-like interaction paradigm. |
| Outcome: | The proposed framework outperforms existing models by +22.5% on the ScreenSpot GUI grounding benchmark. |
Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions (2025.acl-long)
Copied to clipboard
| Challenge: | Experimental results show that multimodal GUI agents are susceptible to environmental distractions. |
| Approach: | They propose a scenario where both user and agent are benign and environment is not malicious . they implement an adversarial environment injection and analyze the approach to improve faithfulness . |
| Outcome: | The proposed approach improves faithfulness of multimodal large language model agents in a graphical user interface environment. |