Papers by Yiqin Wang
Ponder & Press: Advancing Visual GUI Agent towards General Computer Control (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing multimodal large language models (MLLMs) lack visual inputs to ground objects, limiting flexibility across diverse software environments and platforms. |
| Approach: | They propose a divide-and-conquer framework for general computer control that uses only visual inputs to create a purely human-like interaction paradigm. |
| Outcome: | The proposed framework outperforms existing models by +22.5% on the ScreenSpot GUI grounding benchmark. |