Yue Wang
Author directoryOther people with similar names: Yue Wang, Yue Wang, Yue Wang, Yue Wang, Yue Wang
Unverified author pages with similar names: Yue Wang
2026
FedGUI: Benchmarking Federated GUI Agents across Heterogeneous Platforms, Devices, and Operating Systems
WenHao Wang | Haoting Shi | Mengying Yuan | Yiquan Lin | Panrong Tong | Hanzhang Zhou | Guangyi Liu | Pengxiang Zhao | Yue Wang | Siheng Chen
Findings of the Association for Computational Linguistics: ACL 2026
WenHao Wang | Haoting Shi | Mengying Yuan | Yiquan Lin | Panrong Tong | Hanzhang Zhou | Guangyi Liu | Pengxiang Zhao | Yue Wang | Siheng Chen
Findings of the Association for Computational Linguistics: ACL 2026
Training GUI agents with traditional centralized methods faces significant cost and scalability challenges. Federated learning (FL) offers a promising solution, yet its potential is hindered by the lack of benchmarks that capture real-world, cross-platform heterogeneity. To bridge this gap, we introduce FedGUI, the first comprehensive benchmark for developing and evaluating federated GUI agents across mobile, web, and desktop platforms. FedGUI provides a suite of six curated datasets to systematically study four crucial types of heterogeneity: cross-platform, cross-device, cross-OS, and cross-source. Extensive experiments reveal several key insights: First, we show that cross-platform collaboration improves performance, extending prior mobile-only federated learning to diverse GUI environments; Second, we demonstrate the presence of distinct heterogeneity dimensions and identify platform and OS as the most influential factors. FedGUI provides a vital foundation for the community to build more scalable and privacy-preserving GUI agents for real-world deployment. Our code and data are publicly available at https://github.com/wwh0411/FedGUI..
MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments
Quyu Kong | Xu Zhang | Zhenyu Yang | Nolan Gao | Chen Liu | Panrong Tong | Chenglin Cai | Hanzhang Zhou | Jianan Zhang | Liangyu Chen | Zhidan Liu | Steven Hoi | Yue Wang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Quyu Kong | Xu Zhang | Zhenyu Yang | Nolan Gao | Chen Liu | Panrong Tong | Chenglin Cai | Hanzhang Zhou | Jianan Zhang | Liangyu Chen | Zhidan Liu | Steven Hoi | Yue Wang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
While AndroidWorld has become the dominant mobile-use benchmark due to its reproducible environment and deterministic evaluation, recent agents achieving over 90% success rates indicate saturation and motivate the need for greater challenge. In addition, its environment lacks key application categories, such as e-commerce and enterprise communication, and does not reflect realistic mobile-use scenarios characterized by vague user instructions and hybrid tool usage. We introduce MobileWorld, a substantially more challenging benchmark with 201 tasks across 20 applications that reflects real-world usage through long-horizon, cross-application workflows requiring nearly twice as many steps (27.8 vs. 14.3) and featuring significantly more multi-app tasks (62.2% vs. 9.5%) than AndroidWorld. MobileWorld balances production-grade utility and reproducible evaluation using open-source alternatives to industry standards (e.g., Mattermost for Slack), enabling full observability through source code modification and direct database access. Beyond standard GUI manipulation, MobileWorld introduces novel task categories including agent-user interaction and Model Context Protocol (MCP)-augmented tasks for evaluating agents in user-aware, hybrid-tool scenarios. We develop a planner-executor framework with extended action spaces supporting user interactions and MCP calls. Results show a sharp performance drop from AndroidWorld, with the best agentic framework and end-to-end model achieving 51.7% and 20.9% success rates, respectively, highlighting substantial room for future research.
2025
Aria-UI: Visual Grounding for GUI Instructions
Yuhao Yang | Yue Wang | Dongxu Li | Ziyang Luo | Bei Chen | Chao Huang | Junnan Li
Findings of the Association for Computational Linguistics: ACL 2025
Yuhao Yang | Yue Wang | Dongxu Li | Ziyang Luo | Bei Chen | Chao Huang | Junnan Li
Findings of the Association for Computational Linguistics: ACL 2025
Digital agents for automating tasks across different platforms by directly manipulating the GUIs are increasingly important. For these agents, grounding from language instructions to target elements remains a significant challenge due to reliance on HTML or AXTree inputs. In this paper, we introduce Aria-UI, a large multimodal model specifically designed for GUI grounding. Aria-UI adopts a pure-vision approach, eschewing reliance on auxiliary inputs. To adapt to heterogeneous planning instructions, we propose a scalable data pipeline that synthesizes diverse and high-quality instruction samples for grounding. To handle dynamic contexts in task performing, Aria-UI incorporates textual and text-image interleaved action histories, enabling robust context-aware reasoning for grounding. Aria-UI sets new state-of-the-art results across offline and online agent benchmarks, outperforming both vision-only and AXTree-reliant baselines. We release all training data and model checkpoints to foster further research.