Challenge: Existing multimodal large language models struggle with precise localization of small elements.
Approach: They propose a multimodal GUI agent framework that unifies observing, thinking, and acting for precise and interpretable decision-making.
Outcome: The proposed framework unifies observing, thinking, and acting for precise and interpretable decision-making.

Similar Papers

InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection (2026.eacl-long)

Copied to clipboard

Challenge: Existing GUI Agents face challenges in multi-step reasoning and reliance on textual annotations, limiting their effectiveness.
Approach: They propose an MLLM-based GUI Agent with a two-stage supervised fine-tuning pipeline that enhances GUI understanding and grounding.
Outcome: InfiGUIAgent achieves competitive performance on several GUI benchmarks, highlighting the impact of native reasoning skills in enhancing GUI interaction for automation tasks.
Aria-UI: Visual Grounding for GUI Instructions (2025.findings-acl)

Copied to clipboard

Challenge: Using a multimodal model, GUI agents can ground from language instructions to target elements . relying on HTML or AXTree inputs is a challenge for GUI agents .
Approach: They propose a large multimodal model specifically designed for GUI grounding that adopts a pure vision approach instead of auxiliary inputs.
Outcome: The proposed model outperforms vision-only and AXTree-reliant models on offline and online agents.
GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration (2025.emnlp-main)

Copied to clipboard

Challenge: Recent work of GUI action grounding fine-tunes data from pre-trained MLLMs, but data is limited to specific GUI environments.
Approach: They propose to use a GUI-based agent to collect environment-specific data and fine-tune GUI grounding models with the collected data.
Outcome: The proposed model can be extended to other GUI environments to improve performance.
D-Artemis: A Deliberative Cognitive Framework for Mobile GUI Multi-Agents (2026.findings-acl)

Copied to clipboard

Challenge: Graphical User Interface (GUI) agents aim to automate a wide spectrum of human tasks by emulating user interaction.
Approach: They propose a deliberative framework that leverages a fine-grained tip retrieval mechanism to inform its decision-making process.
Outcome: The proposed framework achieves SOTA among open-source general models on AndroidWorld and ScreenSpot-V2 . it leverages a fine-grained, app-specific tip retrieval mechanism to inform its decision-making process .
GUICourse: From General Vision Language Model to Versatile GUI Agent (2025.acl-long)

Copied to clipboard

Challenge: Graphical User Interfaces (GUIs) are a pivotal medium for human-computer interaction.
Approach: They propose a series of datasets for training visual-based GUI agents using general VLMs.
Outcome: The proposed GUICourse datasets show that even a small-sized GUI agent performs better on GUI tasks.
AgentGL: Towards Agentic Graph Learning with LLMs via Reinforcement Learning (2026.acl-long)

Copied to clipboard

Challenge: Existing agentic frameworks treat external information as unstructured text and fail to leverage topological dependencies inherent in real-world data.
Approach: They propose to reframe graph learning as an interleaved process of topology-aware navigation and LLM-based inference.
Outcome: The proposed framework outperforms strong GraphLLMs and GraphRAG benchmarks in multiple LLM backbones.
MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning (2025.acl-long)

Copied to clipboard

Challenge: Existing studies focus on prompting and developing workflows with frozen LLMs.
Approach: They propose a multi-agentic framework for collaborative LLMs with reinforcement learning that leverages multi-gendered frameworks to enhance collaboration.
Outcome: The proposed model improves collaboration performance across multiple datasets with generalization to unseen domains compared to existing models.
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents (2024.acl-long)

Copied to clipboard

Challenge: Existing GUI agents interact with the environment through extracted structured data, which can be notably lengthy (e.g., HTML) and occasionally inaccessible (e-book).
Approach: They propose to enhance SeeClick with GUI grounding pre-training and devise a method to automate curation of GUI ground data.
Outcome: The proposed agent improves ScreenSpot, the first realistic GUI grounding benchmark that encompasses mobile, desktop, and web environments.
You Only Look at Screens: Multimodal Chain-of-Action Agents (2024.findings-acl)

Copied to clipboard

Challenge: Existing approaches to creating autonomous graphical user interfaces rely on external tools and application-specific APIs to interpret the environment.
Approach: They propose a multimodal solution that directly interacts with the user interface without environment parsing.
Outcome: The proposed solution bypasses environment parsing and reliance on application-dependent APIs.
WinSpot: GUI Grounding Benchmark with Multimodal Large Language Models (2025.acl-short)

Copied to clipboard

Challenge: Existing GUI grounding data focuses on web-based elements, leaving a gap in real-world GUI interaction data for non-web applications.
Approach: They propose a framework that leverages Large Language Models to generate large-scale GUI grounding data.
Outcome: The framework validates and refines 5,000 GUI coordinate-instruction pairs and provides high-quality data for training and evaluating visual GUI agents.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations